Electronic device for performing view synthesis, and operating method thereof

By adjusting depth values and selecting appropriate images for synthesis based on camera positioning and shooting direction, the electronic device enhances depth perception and reduces distortions in stereo image display, addressing the challenge of baseline mismatch between internal and external devices.

WO2025146939A1PCT designated stage expired Publication Date: 2025-07-10SAMSUNG ELECTRONICS CO LTD
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
PCT/KR2024/018275
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-05-21
Filing Date
2024-11-19
Publication Date
2025-07-10

AI Technical Summary

Technical Problem

Existing electronic devices struggle to provide a sense of depth similar to real objects when displaying stereo images due to differences in baseline distances between internal cameras and external display devices, leading to image distortions and artifacts.

Method used

An electronic device performs view synthesis by selecting a reference image and a synthesis target image based on the positional relationship of its cameras and the shooting direction, and adjusts depth values to match the baseline of an external display device, minimizing artifacts and enhancing depth perception.

Benefits of technology

The solution allows for accurate depth perception and reduced image distortions, improving the user's experience by ensuring the displayed depth matches real-world depth and maintaining image quality.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure KR2024018275_10072025_PF_FP_ABST
    Figure KR2024018275_10072025_PF_FP_ABST
Patent Text Reader

Abstract

An electronic device for performing view synthesis, and an operating method thereof are provided. The electronic device of the present disclosure: uses a plurality of cameras disposed to be spaced a first baseline apart from each other, so as to photograph an object, thereby acquiring a stereo image including a first image and a second image; on the basis of the positional relationship between the plurality of cameras according to a capturing direction, or positional information of a main object disposed in the image, determines, from among the first image and the second image, a reference image and an image to be synthesized; performs view synthesis on the image to be synthesized that is determined on the basis of the depth information of the object, so as to acquire a synthesized image including the object with a depth value according to a disparity corresponding to a second baseline that is the distance between a left-eye display and a right-eye display of an external device; and can transmit the reference image and the synthesized image to the external device so that the reference image and the synthesized image are displayed by the external device.
Need to check novelty before this filing date? Find Prior Art

Description

Electronic device for performing view synthesis and method of operation thereof

[0001] The present disclosure relates to an electronic device for performing view synthesis and an operating method thereof. Specifically, the present disclosure discloses an electronic device for performing stereo view synthesis to display image content, such as an image or video captured using a stereo camera, through an external device, and an operating method thereof.

[0002] Virtual reality (VR) refers to a specific environment or situation, or the technology itself, created using artificial technology that resembles reality but is not real. Devices that allow users to experience VR include head-mounted displays (HMDs). Among these, see-through HMDs allow users to view the real world and enhance the user experience by augmenting virtual objects onto the real world.

[0003] Augmented reality (AR) is a technology that overlays virtual images onto real-world physical environments or objects. AR devices utilizing AR technology include AR glasses, for example. AR glasses are widely used in everyday life for purposes such as information retrieval, route guidance, and camera photography. Smart glasses, in particular, are often worn as fashion accessories and are primarily used for outdoor activities.

[0004] Users can view image content, such as images or videos taken using a mobile device such as a smart phone, by displaying the image content through a virtual reality device or augmented reality device such as a head-mounted display device. When viewing image content through a virtual reality device, such as a head-mounted display device, the distance between the cameras of the mobile device, i.e., the baseline, is shorter than the baseline between the left-eye display and the right-eye display of the head-mounted display device, so the user cannot feel the depth of the real object from the image content.

[0005] When viewing image content through a virtual reality device or augmented reality device such as a head-mounted display device, in order for the user to feel a sense of depth similar to the depth of a real object, it is necessary to perform stereo view synthesis, which changes the baseline between the cameras according to the distance between the left and right eyes. Stereo view synthesis is an image processing method that synthesizes objects within an image using the depth information of the object. When performing stereo view synthesis, a difference in depth value occurs due to a change in the baseline, and the difference in depth value may cause distortion such as artifacts in the image.

[0006] One aspect of the present disclosure provides an electronic device that performs view synthesis. According to one embodiment of the present disclosure, the electronic device may include a plurality of cameras arranged spaced apart from each other by a first baseline, a communication interface that performs data communication with an external device, a processor including processing circuitry, and a memory that stores one or more instructions. By individually or collectively executing the one or more instructions by at least one processor, the electronic device may acquire a first image and a second image by photographing an object using the plurality of cameras. By individually or collectively executing the one or more instructions by at least one processor, the electronic device may select a reference image from among the first image and the second image based on a positional relationship of the plurality of cameras based on a user's shooting direction when photographing an object or positional information of a main object arranged within the image. By individually or collectively executing the one or more commands by at least one processor, the electronic device can determine the remaining images among the first image and the second image that are not selected as reference images as a synthesis target image. By individually or collectively executing the one or more commands by at least one processor, the electronic device can perform view synthesis on the synthesis target image based on depth information of the object, thereby obtaining a synthesis image including an object having a depth value due to disparity corresponding to a second baseline, which is a distance between a left-eye display and a right-eye display of an external device.By individually or collectively executing one or more of the above commands by at least one processor, the electronic device can control the communication interface to transmit the reference image and the composite image to the external device so that the reference image and the composite image are displayed by the external device.

[0007] One aspect of the present disclosure provides a method for an electronic device to perform view synthesis. According to an embodiment of the present disclosure, an operating method of an electronic device may include a step of capturing an object using a plurality of cameras arranged at a distance from each other by a first baseline, thereby obtaining a stereo image including a first image and a second image. According to an embodiment of the present disclosure, the operating method of an electronic device may include a step of selecting a reference image among the first image and the second image based on a positional relationship of the plurality of cameras based on a shooting direction of a user when capturing an object or positional information of a main object arranged within the image. According to an embodiment of the present disclosure, the operating method of an electronic device may include a step of determining a remaining image among the first image and the second image that is not selected as a reference image as a synthesis target image. An operating method of an electronic device according to an embodiment of the present disclosure may include a step of performing view synthesis on a synthesis target image based on depth information of an object, thereby obtaining a synthesized image including an object having a depth value due to disparity corresponding to a second baseline, which is a distance between a left-eye display and a right-eye display of an external device. An operating method of an electronic device according to an embodiment of the present disclosure may include a step of transmitting a reference image and a synthesized image to an external device such that the reference image and the synthesized image are displayed by the external device.

[0008] One aspect of the present disclosure provides a head mounted display (HMD) device that performs view synthesis. The head mounted display device according to one embodiment of the present disclosure may include a communication interface that pairs with an external device and performs data communication with the external device, a left-eye display and a right-eye display spaced apart by a baseline, at least one processor including a processing circuit, and a memory that stores one or more instructions. The one or more instructions are individually or collectively executed by the at least one processor, thereby controlling the communication interface to receive a stereo image including a first image and a second image acquired by a plurality of cameras included in the external device. By individually or collectively executing the one or more commands by at least one processor, the head-mounted display device can select a reference image among the first image and the second image based on at least one of the positional relationship of a plurality of cameras according to the shooting direction of the external device, the dominant eye of the user, or positional information of a main object placed within the image. By individually or collectively executing the one or more commands by at least one processor, the head-mounted display device can determine the remaining images that are not selected as reference images among the first image and the second image as a synthesis target image.By individually or collectively executing the one or more commands by at least one processor, the head-mounted display device can perform view synthesis on a synthesis target image based on depth information of the object, thereby obtaining a synthesized image including an object having a different depth value at a disparity corresponding to a baseline. By individually or collectively executing the one or more commands by at least one processor, the head-mounted display device can display the reference image and the synthesized image through a left-eye display and a right-eye display.

[0009] The present disclosure can be readily understood by the combination of the following detailed description and the accompanying drawings, wherein reference numerals refer to structural elements.

[0010] Figure 1 is a conceptual diagram illustrating an operation in which an electronic device photographs an object and displays the acquired stereo image through an external device.

[0011] FIG. 2 is a flowchart illustrating a method for an electronic device to perform view synthesis for stereo images according to an embodiment of the present disclosure.

[0012] FIG. 3 is a diagram illustrating an operation of an electronic device according to one embodiment of the present disclosure to perform view synthesis for stereo images.

[0013] FIG. 4 is a block diagram illustrating components of an electronic device according to one embodiment of the present disclosure.

[0014] FIG. 5 is a flowchart illustrating a method for an electronic device according to one embodiment of the present disclosure to select a reference image among stereo images based on a shooting direction and arrangement information of multiple cameras.

[0015] FIG. 6A is a diagram illustrating an embodiment in which an electronic device of the present disclosure selects a right image from among stereo images as a reference image based on a shooting direction and arrangement information of multiple cameras, determines a left image as a synthesis target image, and performs view synthesis for the synthesis target image.

[0016] FIG. 6b is a diagram illustrating an embodiment in which an electronic device of the present disclosure selects a left image from among stereo images as a reference image based on a shooting direction and arrangement information of multiple cameras, determines a right image as a synthesis target image, and performs view synthesis for the synthesis target image.

[0017] FIG. 7 is a flowchart illustrating a method for an electronic device according to one embodiment of the present disclosure to select a reference image based on location information of a main object in an image.

[0018] FIG. 8 is a diagram illustrating an operation of an electronic device according to an embodiment of the present disclosure to recognize a main object from an image using an original image, a segmentation image, and a depth map image.

[0019] FIG. 9A is a diagram illustrating an embodiment in which an electronic device of the present disclosure determines a synthesis target image based on positional information of a major object recognized within an image and performs view synthesis for the synthesis target image.

[0020] FIG. 9b is a diagram illustrating an embodiment in which an electronic device of the present disclosure determines a synthesis target image based on positional information of a major object recognized within an image and performs view synthesis for the synthesis target image.

[0021] Figure 10 is a diagram showing the disparity in the original stereo image and the disparity in the stereo image after view synthesis is performed.

[0022] FIG. 11 is a diagram illustrating a stereo image according to a result of view synthesis performed on an original stereo image by an electronic device according to an embodiment of the present disclosure.

[0023] FIG. 12 is a conceptual diagram illustrating an operation of an electronic device according to an embodiment of the present disclosure to perform view synthesis according to a synthesis target image.

[0024] FIG. 13 is a block diagram illustrating components of a head-mounted display (HMD) device according to one embodiment of the present disclosure.

[0025] FIG. 14 is a flowchart illustrating a method for a head-mounted display device according to one embodiment of the present disclosure to perform view synthesis for stereo images.

[0026] FIG. 15 is a flowchart illustrating an operation method of an electronic device and a head-mounted display device according to one embodiment of the present disclosure.

[0027] FIG. 16 is a flowchart illustrating a method for selecting a reference image for view synthesis based on a user's dominant eye by a head-mounted display device according to one embodiment of the present disclosure.

[0028] FIG. 17 is a diagram illustrating an operation of a head-mounted display device according to one embodiment of the present disclosure to determine a user's gaze by displaying a simulation image.

[0029] The terms used in the embodiments of this specification have been selected from widely used, current terms, taking into account the functions of the present disclosure. However, these terms may vary depending on the intentions of those skilled in the art, precedents, the emergence of new technologies, etc. Furthermore, in certain cases, terms may be arbitrarily selected by the applicant, and in such cases, their meanings will be described in detail in the description of the relevant embodiments. Therefore, the terms used in this specification should not be defined simply as names of terms, but rather based on their meanings and the overall content of the present disclosure.

[0030] Singular expressions may include plural expressions unless the context clearly indicates otherwise. Terms used herein, including technical or scientific terms, have the same meaning as commonly understood by a person of ordinary skill in the art described herein.

[0031] Throughout this disclosure, when a part is said to "include" a component, this does not exclude other components, but rather implies the inclusion of other components, unless otherwise specifically stated. Furthermore, terms such as "part," "module," etc., used herein refer to a unit that processes at least one function or operation, which may be implemented in hardware or software, or a combination of hardware and software.

[0032] As used herein, the expression "configured to" can be used interchangeably with, for example, "suitable for," "having the capacity to," "designed to," "adapted to," "made to," or "capable of." The term "configured to" does not necessarily mean something is "specifically designed to" in terms of hardware. Instead, in some contexts, the expression "a system configured to" can mean that the system is "capable of" in conjunction with other devices or components. For example, the phrase "a processor configured to perform A, B, and C" can mean a dedicated processor for performing the operations (e.g., an embedded processor), or a general-purpose processor (e.g., a CPU or an application processor) that can perform the operations by executing one or more software programs stored in memory.

[0033] Additionally, when a component is referred to as being "connected" or "connected" to another component in the present disclosure, it should be understood that the component may be directly connected or connected to the other component, but may also be connected or connected via another component in between, unless otherwise specifically stated.

[0034] All functions or operations described in this disclosure may be processed by a single processor or a combination of multiple processors. The single processor or the combination of multiple processors may include circuitry that performs processing, such as an Application Processor (AP), a Communication Processor (CP), a Graphical Processing Unit (GPU), a Neural Processing Unit (NPU), a Microprocessor Unit (MPU), a System on Chip (SoC), or an Integrated Chip (IC).

[0035] It should be understood that the blocks and combinations of flowcharts illustrated in this disclosure can be implemented by one or more computer programs containing computer-executable instructions. The one or more computer programs may be stored entirely in a single memory, or may be divided and stored across multiple different memories.

[0036] In the present disclosure, 'view synthesis' means image processing that changes the depth value of an object and synthesizes it so that the object has the same depth value as when photographed from a desired camera position using an image (e.g., an RGB image) and depth information of an object (e.g., a depth map image). In one embodiment of the present disclosure, view synthesis may mean stereo view synthesis that changes the depth value of an object in an image by adjusting the disparity between matched pixels of each of a plurality of images constituting a stereo image, and synthesizes the object with the changed depth value with a background.

[0037] In the present disclosure, 'stereo view synthesis' refers to image processing that shifts an object included in one of the left and right images that constitute a stereo image based on the depth value of the object, and synthesizes the shifted object with a background. In one embodiment of the present disclosure, stereo view synthesis may be performed by selecting one of the left and right images as a reference image, and targeting the remaining images that are not selected as reference images.

[0038] Below, embodiments of the present disclosure are described in detail with reference to the attached drawings so that those skilled in the art can easily implement the present disclosure. However, the present disclosure may be implemented in various different forms and is not limited to the embodiments described herein.

[0039] Hereinafter, embodiments of the present disclosure will be described in detail with reference to the drawings.

[0040] FIG. 1 is a conceptual diagram illustrating an operation in which an electronic device (100) photographs an object and displays the acquired stereo images (i1, i2) through an external device (200).

[0041] The electronic device (100) may be a smart phone or tablet PC including multiple cameras (111, 112). However, it is not limited to the illustrated embodiment, and the electronic device (100) may be implemented as a mobile device such as, for example, a laptop computer, a digital camera, an e-book terminal, a digital broadcasting terminal, a PDA (Personal Digital Assistant), a PMP (Portable Multimedia Player), a navigation device, or an MP3 player.

[0042] Referring to FIG. 1, the electronic device (100) can capture an object (10) through a plurality of cameras (111, 112) to obtain a plurality of images (i1, i2). In one embodiment of the present disclosure, the plurality of cameras (111, 112) can be implemented as stereo cameras configured to capture stereo images. The plurality of cameras (111, 112) can be arranged on the electronic device (100) to be spaced apart from each other by a first baseline (b1). The first baseline (b1) can be, for example, 1.5 centimeters (cm), but is not limited thereto. The electronic device (100) can capture an object (10) using the first camera (111) to obtain a first image (i1), and can capture an object (10) using the second camera (112) to obtain a second image (i2). The electronic device (100) can transmit the acquired first image (i1) and second image (i2) to an external device (200).

[0043] The external device (200) may be configured as an augmented reality (AR) device or a virtual reality (VR) device. The external device (200) may be implemented as, for example, a head-mounted display device. In the present disclosure, a 'head-mounted display (HMD) device' is a device worn on the user's head and provides a virtual reality experience to the user. However, the present invention is not limited thereto, and the external device (200) may be implemented as, for example, an augmented reality helmet (Augmented Reality Helmet) that provides an augmented reality experience, a face-mounted display (FMD) device worn on the user's face, or augmented reality glasses in the shape of glasses.

[0044] Hereinafter, it will be described that the external device (200) is implemented as a head-mounted display device.

[0045] A head-mounted display device (200) can display a first image (i1) and a second image (i2) received from an electronic device (100). The head-mounted display device (200) includes a left-eye display (220L) and a right-eye display (220R), and can display the first image (i1) through the left-eye display (220L) and display the second image (i2) through the right-eye display (220R). In this case, the first image (i1) can be a left-eye image, and the second image (i2) can be a right-eye image.

[0046] In a head-mounted display device (200), the left-eye display (220L) and the right-eye display (220R) may be spaced apart by a second baseline (b2). In one embodiment of the present disclosure, the second baseline (b2) may be an average value of the distance between a person's left and right eyes. For example, the second baseline (b2) may be 6.5 centimeters (cm). However, the present disclosure is not limited thereto.

[0047] A first baseline (b1), which is a distance between a plurality of cameras (111, 112) arranged in an electronic device (100), is relatively smaller than a second baseline (b2), which is a distance between a plurality of displays (220L, 220R) of a head-mounted display device (200). When the electronic device (100) does not synthesize the first image (i1) and the second image (i2) but transmits them to the head-mounted display device (200) in an original image state, the depth value of an object (10) displayed in stereo by the left-eye image and the right-eye image may change due to a change in disparity. As illustrated in FIG. 1, the object (10') displayed by the head-mounted display device (200) through the left-eye display (220L) and the right-eye display (220R) is different from the depth values ​​of the first image (i1) and the second image (i2) captured by the electronic device (100), and the user may feel a weak sense of depth.

[0048] When an electronic device (100) performs view synthesis so that the depth values ​​of objects in a first image (i1) and a second image (i2) have depth values ​​according to a parallax corresponding to a second baseline (b2), and transmits the synthesized image to a head-mounted display device (200), an object (10) displayed by the head-mounted display device (200) may have a depth value that is the same as or similar to the depth value of an object (10) photographed by the electronic device (100). However, image distortion, such as artifacts due to view synthesis, may occur.

[0049] Referring to FIG. 1, the electronic device (100) may select the first image (i1), which is the left image, among the first image (i1) and the second image (i2) constituting the stereo image, as a reference image, and select the second image (i2), which is the right image, as a synthesis target image. The electronic device (100) may perform view synthesis to shift an object (10R) in the selected synthesis target image so that it has a disparity corresponding to the second baseline (b2) in relation to an object (10L) in the reference image. The synthesized image (isynthesized) obtained through view synthesis may include an object (10R') moved to the left during the view synthesis process and an artifact (10a) caused by the moved object (10R'). The object (10R') of the synthesized image (isynthesized) may be an original image, which is a reference image (i ref ) can be synthesized by moving to the left so as to have a depth value corresponding to the second baseline (b2) due to the disparity formed in the relationship with the object (10L). In the view synthesis process, an artifact (10a) may be generated as an area where image information does not exist is arbitrarily synthesized due to the movement of the object (10R).

[0050] FIG. 2 is a flowchart illustrating a method by which an electronic device (100) according to one embodiment of the present disclosure performs view synthesis for stereo images.

[0051] FIG. 3 is a diagram illustrating an operation of an electronic device (100) according to one embodiment of the present disclosure to perform view synthesis for a stereo image.

[0052] Hereinafter, the function and / or operation of an electronic device (100) according to one embodiment of the present disclosure will be described in detail with reference to FIGS. 2 and 3 together.

[0053] Referring to FIG. 2, in step S210, the electronic device (100) acquires a stereo image through a plurality of cameras spaced apart by a first baseline.

[0054] Referring to the embodiment illustrated in FIG. 3, the electronic device (100) includes a plurality of cameras (111, 112), and can obtain a stereo image including a first image (i1) and a second image (i2) using the plurality of cameras (111, 112). The number and arrangement of the plurality of cameras (111, 112) are not limited to the embodiment illustrated in FIG. 3. In one embodiment of the present disclosure, the electronic device (100) may include a single camera or three or more multi-cameras, and the plurality of cameras (111, 112) may also be arranged in a manner other than in a row.

[0055] In one embodiment of the present disclosure, the plurality of cameras (111, 112) may be stereo cameras configured to capture stereo images. The plurality of cameras (111, 112) may be arranged on the electronic device (100) spaced apart from each other by a first baseline (b1). The first baseline (b1) may be, for example, 1.5 centimeters (cm), but is not limited thereto. Referring to operation ① of FIG. 3, the electronic device (100) may capture an object using the first camera (111) to obtain a first image (i1), and may capture an object using the second camera (112) to obtain a second image (i2). In the embodiment illustrated in FIG. 3, the first image (i1) may be a left image positioned on the left, and the second image (i2) may be a right image positioned on the right, depending on the shooting direction determined by the direction in which the user holds the electronic device (100) when capturing an object.

[0056] Referring back to FIG. 2, in step S220, the electronic device (100) selects a reference image among stereo images based on the positional relationship of a plurality of cameras according to the user's shooting direction or the positional information of a main object within the image. In one embodiment of the present disclosure, the electronic device (100) may recognize the shooting direction according to the posture in which the electronic device (100) is held by the user during shooting, and select a reference image among the first image and the second image based on the recognized shooting direction and the positional information of the plurality of cameras arranged on the electronic device (100).

[0057] In step S230, the electronic device (100) determines an image that is not selected as a reference image among the multiple images included in the stereo image as a synthesis target image.

[0058] Referring to the embodiment illustrated in FIG. 3, the electronic device (100) includes a gravity sensor that obtains a measurement value regarding the direction of gravity, and can recognize the shooting direction regarding whether the electronic device (100) is held vertically or horizontally by the user when taking a picture using the gravity sensor. When taking a picture while being held horizontally, the electronic device (100) can recognize the shooting direction regarding whether the picture is taken in a direction in which the plurality of cameras (111, 112) are positioned to the left or in a direction in which the plurality of cameras (111, 112) are positioned to the right based on the measurement value by the gravity sensor. The electronic device (100) can select a reference image, which is an original image that serves as a basis for view synthesis, from among the first image (i1) and the second image (i2) acquired by the plurality of cameras (111, 112) based on the recognized shooting direction and the arrangement information regarding the direction in which the plurality of cameras (111, 112) are arranged on the electronic device (100).

[0059] In one embodiment of the present disclosure, the electronic device (100) identifies a camera that is farthest from the center line (l) of the electronic device (100) among the plurality of cameras (111, 112) based on the positional relationship of the plurality of cameras (111, 112), and captures an image captured by the identified camera among the first image (i1) and the second image (i2) as a reference image (i ref ) can be selected as. Referring to operation ② of FIG. 3, the electronic device (100) recognizes the shooting direction by the user based on the measurement value acquired by the gravity sensor, and can recognize that the plurality of cameras (111, 112) are arranged to the right with respect to the center line (l) of the electronic device (100) based on the recognized shooting direction and the position information of the plurality of cameras (111, 112). In this case, the electronic device (100) identifies the second camera (112) that is arranged farthest from the center line (l) among the plurality of cameras (111, 112), and sets the second image (i2) captured by the identified second camera (112) as the reference image (i ref ) can be selected. The electronic device (100) can identify the first camera (111) positioned adjacent to the center line (l) among the plurality of cameras (111, 112) as a source camera, and determine the first image (i1) captured by the first camera (111) identified as the source camera as a synthesis target image.

[0060] In the opposite case to the embodiment illustrated in FIG. 3, that is, when multiple cameras (111, 112) are arranged to be offset to the left with respect to the center line (l) according to the shooting direction by the user, the electronic device (100) identifies the first camera (111) that is arranged to be farthest from the center line (l) among the multiple cameras (111, 112), and sets the first image (i1) captured by the first camera (111) as the reference image (i ref) can be selected. The electronic device (100) selects a reference image (i) among the first image (i1) and the second image (i2). ref ) can be determined as the second image (i2), which is the remaining image that was not selected as the target image for synthesis.

[0061] In one embodiment of the present disclosure, the electronic device (100) recognizes a main object from a first image (i1) and a second image (i2), and recognizes a reference image (i) based on the location information of the main object within the image. ref ) and a target image for synthesis can be determined. Referring to operation ② of FIG. 3, the electronic device (100) performs image segmentation to recognize at least one object from at least one of the first image (i1) and the second image (i2), and determines the main object (20L, 20R) based on at least one of the position, direction, size, depth, and movement of the recognized at least one object. The electronic device (100) determines the reference image (i) based on the location information about where the main object (20L, 20R) is located in the first image (i1) and the second image (i2). ref ) and a synthesis target image can be determined. For example, if the main object (20L, 20R) is positioned to the left in the image, the electronic device (100) can select the first image (i1), which is the left image, as the synthesis target image. In the embodiment illustrated in FIG. 3, since the main object (20L, 20R) is positioned to the left in both the first image (i1) and the second image (i2), the electronic device (100) can select the first image (i1), which is the left image, as the synthesis target image for performing view synthesis.

[0062] As an opposite example, if the main object (20L, 20R) is positioned to the right within the image, the electronic device (100) may select the second image (i2), which is the right image, as the target image for synthesis.

[0063] Referring back to FIG. 2, in step S240, the electronic device (100) performs view synthesis on a synthesis target image based on the depth information of the object, thereby obtaining a synthesized image including an object having a depth value due to disparity corresponding to a second baseline, which is a distance between displays of an external device. In the present disclosure, 'view synthesis' means image processing that synthesizes an image by changing the depth value of an object so that the object has the same depth value as when photographed at a desired camera position using depth information of an image and an object. In one embodiment of the present disclosure, view synthesis may mean stereo view synthesis that changes the depth value of an object in an image by adjusting the disparity between matched pixels of each of a plurality of images constituting a stereo image, and synthesizes the object with the changed depth value with a background.

[0064] The electronic device (100) can acquire depth information of an object through stereo imaging using a left image and a right image, and perform view synthesis using the acquired depth information. In one embodiment of the present disclosure, the electronic device (100) can extract feature points from the left image and the right image, match the feature points extracted from each of the left image and the right image, calculate disparity according to the distance between pixels constituting the matched feature points, and acquire a depth value of the object using the calculated disparity. However, the present invention is not limited thereto, and the electronic device (100) can include a depth sensor configured to measure a depth value, such as a Time-of-Flight (ToF) sensor or a Light wave Detection and Ranging (LiDAR) sensor, and can also acquire depth information of the object using the depth sensor.

[0065] Referring to operation ③ of FIG. 3 together, the electronic device (100) can perform view synthesis so that the first image (i1) selected as the synthesis target image has a depth value due to disparity as if it were captured at a virtual camera position (P) spaced apart from the second camera (112) by a second baseline (b2). The second baseline (b2) is an average value of the distance between the user's left and right eyes, and may be, for example, 6.5 cm. In one embodiment of the present disclosure, the electronic device (100) can perform view synthesis so that the object (20L) of the first image (i1) selected as the synthesis target image and the reference image (i ref ) can perform view synthesis by shifting the object (20L) of the synthesis target image to the right so that the object (20R) in the image has a depth value according to the disparity corresponding to the second baseline (b2), and synthesizing the shifted object (20L') with the background.

[0066] Unlike the embodiment illustrated in FIG. 3, when the second image (i2) is determined as the synthesis target image, the electronic device (100) can perform view synthesis by shifting the object (20R) of the synthesis target image to the left so that the object (20R) of the second image (i2) and the object (20L) of the reference image (in this case, the 'first image (i1)') have depth values ​​according to the disparity corresponding to the second baseline (b2), and synthesizing the shifted object with the background.

[0067] Referring again to FIG. 2, in step S250, the electronic device (100) transmits the reference image and the synthesized image to the external device. Referring also to operation ④ of FIG. 3, among the stereo images including the first image (i1) and the second image (i2), the first image (i1) is determined as the synthesis target image and becomes the synthesized image (isynthesized) through view synthesis, and the second image (i2) is the reference image (i ref ) is an original image that has not been synthesized. The electronic device (100) is a reference image (i ref ) and a synthesized image (isynthesized) can be transmitted to an external device (200). The external device (200) may be configured as an augmented reality (AR) device or a virtual reality (VR) device. The external device (200) may be implemented as, for example, a head-mounted display device.

[0068] Referring to operation ⑤ of FIG. 3, the external device (200) is implemented as a head mounted display (HMD) device, and the head mounted display device receives a reference image (i) from the electronic device (100). ref) and a synthesized image (isynthesized). In one embodiment of the present disclosure, the head mounted display device includes a left-eye display (220L) and a right-eye display (220R), and displays a reference image (i) through the left-eye display (220L). ref ) can be displayed, and a synthetic image (isynthesized) can be displayed through the right display (220R). However, the present disclosure is not limited to that shown in FIG. 3. In one embodiment of the present disclosure, the first image (i1), which is the left image, can be displayed as a reference image (i ref ) is selected, and the second image (i2), which is the right image, is determined as the synthesis target image, and the second image (i2) becomes the synthesis image (isynthesized) through view synthesis, the reference image (i ref ) can be displayed through the left eye display (220L) of the head-mounted display device, and a synthesized image (isynthesized) can be displayed through the right eye display (220R).

[0069] A user can view image content by taking an image or video using an electronic device (100) such as a smart phone and displaying the taken image or video through an external device such as a head-mounted display device. Since the first baseline (b1), which is the distance between the cameras of the electronic device (100), is shorter than the second baseline (b2), which is the distance between the binocular displays (220L, 220R) of an external device, for example, a head-mounted display device, when viewing image content through an external device, the user cannot feel the depth of a real object from the image content. In order for the user to feel a depth similar to the depth of a real object when viewing image content through an external device such as a head-mounted display device, it is necessary to perform stereo view synthesis, which changes the first baseline (b1) between the cameras according to the second baseline (b2), which is the distance between the left and right eyes. Referring to FIG. 1, when performing stereo view synthesis, a difference in depth value occurs due to a change in the baseline, and the difference in depth value may cause distortion in the image, such as an artifact (10a, see FIG. 1). If the image is distorted by the artifact (10a) due to view synthesis, the quality of the image may deteriorate, and the satisfaction of a user viewing image content through a head-mounted display device (200, see FIG. 1) may decrease.

[0070] The present disclosure provides an electronic device (100) and an operating method thereof for synthesizing stereo images so as to minimize the occurrence of artifacts and to enable a sense of depth identical to or similar to the depth value of a real object without deterioration of image quality when performing view synthesis by changing a first baseline (b1), which is a distance between a plurality of cameras (111, 112), to a second baseline (b2) corresponding to the distance between the left and right eyes of a user (e.g., 6.5 cm).

[0071] The electronic device (100) according to the embodiment illustrated in FIGS. 2 and 3 selects a reference image and a synthesis target image from among stereo images including a first image (i1) and a second image (i2) based on the positional relationship of a plurality of cameras (111, 112) according to the shooting direction of the user when shooting an object or the positional information of a main object placed in the image, and performs view synthesis on the synthesis target image so that the object included in the selected synthesis target image based on the depth information of the object has a depth value according to the disparity corresponding to the second baseline (b2), thereby obtaining a synthesized image, and transmits the reference image and the synthesis image to an external device (e.g., a 'head-mounted display device (200)'), so that when the reference image and the synthesis image are displayed through the left-eye display (220L) and the right-eye display (220R) of the head-mounted display device (200) spaced apart by the second baseline (b2) corresponding to the distance between the left and right eyes, the user can see the reality. It can provide a sense of depth identical to that of an object. In addition, the electronic device (100) according to one embodiment of the present disclosure can minimize the occurrence of artifacts due to view synthesis and provide a technical effect of improving image quality by selecting a synthesis target image and synthesizing the selected synthesis target image.

[0072] FIG. 4 is a block diagram illustrating components of an electronic device (100) according to one embodiment of the present disclosure.

[0073] The electronic device (100) may be implemented as a mobile device such as a smart phone, tablet PC, laptop computer, digital camera, e-book terminal, digital broadcasting terminal, PDA (Personal Digital Assistants), PMP (Portable Multimedia Player), navigation, or MP3 player including a camera (110).

[0074] Referring to FIG. 4, the electronic device (100) may include a camera (110), a sensor (120), a processor (130), a memory (140), and a communication interface (150). The camera (110), the sensor (120), the processor (130), the memory (140), and the communication interface (150) may each be electrically and / or physically connected to each other. In FIG. 4, only essential components for explaining the function and / or operation of the electronic device (100) are illustrated, and the components included in the electronic device (100) are not limited as illustrated in FIG. 4. In one embodiment of the present disclosure, the electronic device (100) may further include a battery that supplies driving power to the camera (110), the sensor (120), the processor (130), and the communication interface (150). In one embodiment of the present disclosure, the electronic device (100) may further include a display unit that displays an image or video acquired through the camera (110).

[0075] The camera (110) is configured to capture an image by photographing an object. The camera (110) may include a lens module, an image sensor, and an image processing module. The camera (110) may capture a still image or a moving image of the object by the image sensor (e.g., CMOS or CCD). The moving image may include a plurality of image frames continuously captured by photographing the object through the camera (110). The image processing module may store a still image composed of a single image frame captured by the image sensor or moving image data composed of a plurality of image frames in the image data storage (146) within the memory (140).

[0076] In one embodiment of the present disclosure, the camera (110) may be implemented in a small form factor so that it can be mounted on a mobile device, and may be implemented as a lightweight RGB camera that consumes low power.

[0077] The camera (110) may include two or more cameras. The camera (110) may be implemented as a plurality of stereo cameras. In the embodiment illustrated in FIG. 4, the camera (110) may include a first camera (111) and a second camera (112). However, the present disclosure is not limited thereto, and in one embodiment of the present disclosure, the electronic device (100) may include three or more cameras. The first camera (111) and the second camera (112) may be arranged spaced apart from each other by a first baseline.

[0078] The electronic device (100) can obtain a first image by photographing an object using a first camera (111), and can obtain a second image by photographing an object using a second camera (112). The first image and the second image may be stereo images.

[0079] The sensor (120) is configured to sense movement, rotation, or speed change of the electronic device (100) to obtain a measurement value. In one embodiment of the present disclosure, the sensor (120) may include a gravity sensor.

[0080] A gravity sensor is a sensor configured to detect the Earth's gravity and obtain measurements regarding the direction and degree of gravity. In one embodiment of the present disclosure, the gravity sensor may obtain measurements regarding the direction of gravity and provide the measurements to the processor (130). The processor (130) may determine whether the electronic device (100) is placed horizontally or vertically based on the measurements obtained from the gravity sensor. If the electronic device (100) is placed horizontally, the processor (130) may determine whether the electronic device (100) is placed leftward or rightward based on the measurements obtained from the gravity sensor.

[0081] In one embodiment of the present disclosure, the sensor (120) may further include an IMU sensor or a position sensor.

[0082] An IMU (Inertial Measurement Unit) sensor is a sensor configured to measure the moving speed, direction, angle, and gravitational acceleration of an electronic device (100) through a combination of an accelerometer, a gyroscope, and a magnetometer. The processor (130) can obtain 6 DoF (6 Degree of Freedom) measurement values ​​including 3D position coordinate values ​​(x-axis, y-axis, and z-axis coordinate values) and 3-axis angular velocity values ​​(roll, yaw, pitch) of the electronic device (100) using the IMU sensor.

[0083] The location sensor is a sensor configured to obtain location information of an electronic device (100), and may be implemented as, for example, a GPS (Global Positioning System) sensor.

[0084] In one embodiment of the present disclosure, the sensor (120) may include a depth sensor. The 'depth sensor' is a sensor configured to acquire depth information by measuring the depth values ​​of objects, and may be implemented as, for example, a Time-of-Flight (ToF) sensor or a Light Wave Detection and Ranging (LiDAR) sensor. However, the depth sensor is not limited to the examples described above. The processor (130) may measure the depth values ​​of an object through the depth sensor and acquire a depth map image.

[0085] The processor (130) can execute one or more instructions of a program stored in the memory (140). The processor (130) may be composed of hardware components that perform arithmetic, logic, input / output operations, and image processing. Although the processor (130) is illustrated as a single element in FIG. 4, it is not limited thereto. In one embodiment of the present disclosure, the processor (130) may be composed of one or more elements.

[0086] The processor (130) may include various processing circuits and / or multiple processors. For example, the term "processor" as used herein, including the claims, may include various processing circuits, including at least one processor. One or more processors in at least one processor may be configured to perform various functions described herein, individually and / or collectively, in a distributed fashion. As used herein, "processor," "at least one processor," and "one or more processors" may be configured to perform various functions. However, these terms encompass, without limitation, situations where one processor performs some of the functions and other processor(s) perform other parts of the functions, and situations where a single processor may perform all of the functions. Furthermore, at least one processor may include a combination of processors that perform various functions of the disclosed functions in a distributed manner. At least one processor may execute program instructions to achieve or perform various functions.

[0087] One or more processors included in the processor (130) may be circuitry such as a system on chip (SoC), an integrated circuit (IC), etc. The processor (130) may be implemented as a general-purpose processor such as a central processing unit (CPU), an application processor (AP), a digital signal processor (DSP), a graphics-only processor such as a graphics processing unit (GPU), a vision processing unit (VPU), or an artificial intelligence-only processor such as a neural processing unit (NPU), for example. The processor (130) may be controlled to process input data according to a predefined operation rule or artificial intelligence model. Alternatively, when the processor (130) is an artificial intelligence-only processor, the artificial intelligence-only processor may be designed with a hardware structure specialized for processing a specific artificial intelligence model.

[0088] The memory (140) may be configured as at least one type of storage medium, for example, a flash memory type, a hard disk type, a multimedia card micro type, a card type memory (e.g., SD or XD memory, etc.), a RAM (Random Access Memory), a SRAM (Static Random Access Memory), a ROM (Read-Only Memory), an EEPROM (Electrically Erasable Programmable Read-Only Memory), a PROM (Programmable Read-Only Memory), or an optical disk.

[0089] The memory (140) may store instructions related to functions and / or operations for the electronic device (100) to select a reference image among stereo images, determine a synthesis target image, and perform view synthesis for the synthesis target image. In one embodiment of the present disclosure, the memory (140) may store at least one of instructions, an algorithm, a data structure, a program code, and an application program that can be read by the processor (130). The instructions, algorithms, data structures, and program codes stored in the memory (140) may be implemented in a programming or scripting language such as, for example, C, C++, Java, or an assembler.

[0090] The memory (140) may store instructions, algorithms, data structures, or program codes related to the reference image selection module (142) and the view synthesis module (144). A 'module' included in the memory (140) refers to a unit that processes a function or operation performed by the processor (130), and this may be implemented as software such as instructions, algorithms, data structures, or program codes. In one embodiment of the present disclosure, the memory (140) may also include image data storage (146).

[0091] The processor (130) may be implemented by executing instructions or program codes stored in the memory (140). Hereinafter, the functions and / or operations performed by the processor (130) by executing instructions or program codes of each of the plurality of modules stored in the memory (140) will be described in detail.

[0092] The reference image selection module (142) is configured with commands or program codes for executing a function and / or operation of selecting a reference image for view synthesis among a plurality of images constituting a stereo image. In one embodiment of the present disclosure, the reference image selection module (142) may select a reference image among stereo images based on positional relationship information of a plurality of cameras according to a user's shooting direction or positional information of a main object within the image. The processor (130) may select a reference image among stereo images captured by a plurality of cameras (111, 112) by executing the commands or program codes of the reference image selection module (142). However, the present invention is not limited thereto, and the processor (130) may also select a reference image among stereo images acquired in advance and stored in the image data storage (146).

[0093] The processor (130) can recognize the shooting direction, based on the measurement value of the gravity direction acquired from the gravity sensor, regarding whether the electronic device (100) is held vertically or horizontally by the user when capturing an image. When capturing an image while being held horizontally, the processor (130) can recognize the shooting direction, based on the measurement value by the gravity sensor, regarding whether the shooting is performed in a direction in which the plurality of cameras (111, 112) are positioned to the left or in a direction in which the plurality of cameras (111, 112) are positioned to the right. The processor (130) can select a reference image, which is an original image that serves as a basis for view synthesis, from among the first image and the second image acquired by the plurality of cameras (111, 112), based on the recognized shooting direction and the arrangement information regarding the direction in which the plurality of cameras (111, 112) are positioned on the electronic device (100). The processor (130) can determine the remaining images that are not selected as reference images among the first and second images as the target images for synthesis. A specific embodiment in which the processor (130) recognizes the shooting direction based on the measurement value acquired by the gravity sensor, selects the reference image among the stereo images based on the positional relationship of the plurality of cameras (111, 112) according to the shooting direction and the arrangement relationship of the plurality of cameras (111, 112), and determines the target image for synthesis will be described in detail with reference to FIGS. 5, 6A, and 6B.

[0094] The processor (130) may recognize a main object from at least one of the stereo images, and select a reference image based on position information of the main object within the image. In one embodiment of the present disclosure, the processor (130) may perform image segmentation to recognize at least one object from at least one of the first image and the second image included in the stereo image, and determine the main object based on at least one of the position, direction, size, depth, and movement of the recognized at least one object. The processor (130) may determine the reference image and the synthesis target image based on position information regarding where the main object is located within the first image and the second image. In one embodiment of the present disclosure, when the main object is located in a first direction with respect to the center of the image, the processor (130) may select an image in which the main object is located closest to the center of the image among the first image and the second image as the reference image, and determine the remaining images among the first image and the second image that are not selected as the reference image as the synthesis target image. For example, if the main object is positioned to the left within the image, the processor (130) may select an image in which the main object is positioned closest to the center of the image among the first image and the second image as a reference image, and determine the first image, which is the left image, as the synthesis target image. For the opposite example, if the main object is positioned to the right within the image, the processor (130) may determine the second image, which is the right image, as the synthesis target image. Specific embodiments in which the processor (130) recognizes a main object within an image and determines a reference image and a synthesis target image based on position information of the main object will be described in detail with reference to FIGS. 7, 8, 9A, and 9B.

[0095] The view synthesis module (144) is configured with commands or program codes for executing a function and / or operation of performing view synthesis on a synthesis target image based on depth information of an object to obtain a synthesized image. In the present disclosure, 'view synthesis' means image processing that changes the depth value of an object and synthesizes it by using an image (e.g., an RGB image) and depth information of an object (e.g., a depth map image) so that the object has the same depth value as when photographed at a desired camera position. In one embodiment of the present disclosure, view synthesis may mean stereo view synthesis that changes the depth value of an object in an image by adjusting the disparity between the matched pixels of each of a plurality of images constituting a stereo image, and synthesizes the object with the changed depth value with the background. The processor (130) can perform view synthesis on a synthesis target image by executing commands or program codes of the view synthesis module (144), thereby obtaining a synthesized image.

[0096] The first camera (111) and the second camera (112) are arranged to be spaced apart by a first baseline, so that the stereo images acquired through the first camera (111) and the second camera (112) can have a disparity corresponding to the first baseline. The processor (130) can acquire a composite image including an object having a depth value due to a disparity corresponding to the second baseline, which is a distance between displays of an external device, by performing view synthesis on the composite target image based on the depth information of the object. The 'second baseline' is an average value of the distance between the left and right eyes of the user, and can be, for example, 6.5 cm. In one embodiment of the present disclosure, the processor (130) can perform view synthesis by shifting an object of the composite target image so that an object included in the first image selected as the composite target image and an object included in the reference image have a depth value due to a disparity corresponding to the second baseline, and synthesizing the shifted object with the background. The parallax that changes as an object moves by performing view composition is described in detail with reference to Figure 10.

[0097] When synthesizing views, the reference image may be maintained as the original image without being synthesized. However, in one embodiment of the present disclosure, the processor (130) may perform view synthesis not only on the synthesis target image but also on the reference image. In this case, the processor (130) may perform view synthesis by moving an object included in the synthesis target image in a first direction so that the object has a depth value according to parallax corresponding to the second baseline, and moving an object included in the reference image in a second direction opposite to the first direction, thereby obtaining a second synthesized image. Various view synthesis methods depending on whether synthesis is performed on the reference image will be described in detail in FIG. 12.

[0098] Image data storage (146) is a storage device in memory (140) that stores stereo images acquired by being photographed by multiple cameras (111, 112) or stereo images acquired in advance through a network such as a website. Image data storage (146) may be configured as non-volatile memory. Non-volatile memory refers to a storage medium that stores and maintains information even when power is not supplied and can use the stored information again when power is supplied. Non-volatile memory may include at least one of, for example, flash memory, a hard disk, an SSD (Solid State Drive), a multimedia card micro type, a card type memory (for example, SD or XD memory, etc.), a ROM (Read Only Memory; ROM), a magnetic memory, a magnetic disk, and an optical disk.

[0099] Although the image data storage (146) is illustrated in FIG. 4 as a component included within the memory (140), the present disclosure is not limited to the configuration illustrated in the drawing. In one embodiment of the present disclosure, the image data storage (146) may be configured as a database within the electronic device (100), which is a separate component from the memory (140).

[0100] However, the present disclosure is not limited thereto, and in one embodiment of the present disclosure, the image data storage (146) may be implemented as a web storage or cloud server that is accessible via a network and performs a storage function. In this case, the electronic device (100) may communicate with the web storage or cloud server through a communication interface (150) and perform data transmission and reception. The processor (130) may receive stereo images from the web storage or cloud server.

[0101] The processor (130) can control the communication interface (150) to transmit the reference image and the composite image to an external device. The communication interface (150) is a hardware device configured to perform data communication with an external device and / or a server. The communication interface (150) can be configured as a device that performs data communication with an external device or a server using at least one of data communication methods including, for example, wireless LAN, Wi-Fi, Wi-Fi Direct, Bluetooth, Bluetooth Low Energy (BLE), infrared Data Association (IrDA), Near Field Communication (NFC), Wireless Broadband Internet (Wibro), World Interoperability for Microwave Access (WiMAX), Shared Wireless Access Protocol (SWAP), Wireless Gigabit Alliance (WiGig), and RF communication. The external device is a device that is connected to the electronic device (100) through a communication network, and may be, for example, a head-mounted display device.

[0102] In one embodiment of the present disclosure, a communication interface (150) can be connected to a head-mounted display device (200, see FIG. 3) under the control of a processor (130) and transmit and receive data. The communication interface (150) can pair with the head-mounted display device (200) via a short-range wireless communication network, for example, Bluetooth, BLE, or Wi-Fi Direct, and transmit a reference image and a composite image to the head-mounted display device (200). The head-mounted display device (200) can display the reference image and the composite image via a left-eye display (220L, see FIG. 3) and a right-eye display (220R, see FIG. 3).

[0103] FIG. 5 is a flowchart illustrating a method for an electronic device (100) according to one embodiment of the present disclosure to select a reference image among stereo images based on a shooting direction and arrangement information of multiple cameras.

[0104] Steps S510 to S540 illustrated in FIG. 5 are operations that specify the operation of step S220 of FIG. 2. Step S510 illustrated in FIG. 5 may be performed after the operation of step S210 of FIG. 2 is performed. After the operation of step S540 illustrated in FIG. 5 is performed, the operation of step S230 of FIG. 2 may be performed.

[0105] FIG. 6a shows an electronic device (100) of the present disclosure that captures a right image (i) among stereo images based on the shooting direction and arrangement information of multiple cameras (111, 112). R ) as a reference image (i ref ) and select the left image (i L ) is determined as a synthesis target image, and a diagram illustrating an embodiment of performing view synthesis on the synthesis target image.

[0106] Hereinafter, the function and / or operation of the electronic device (100) will be described in detail with reference to FIG. 5 and FIG. 6a together.

[0107] In step S510 of FIG. 5, the electronic device (100) recognizes the shooting direction based on the measurement value acquired by the gravity sensor. In one embodiment of the present disclosure, the electronic device (100) may include a gravity sensor configured to acquire a measurement value of the gravity direction. The electronic device (100) may recognize the shooting direction by using the gravity sensor, regarding whether the electronic device (100) is held vertically or horizontally by the user when taking a picture. Referring also to operation ① of FIG. 6A, when the electronic device (100) is held by the user while being rotated 90° in the left direction and taking a picture, the processor (130, see FIG. 4) of the electronic device (100) recognizes the shooting direction based on the measurement value by the gravity sensor, and can identify that a plurality of cameras (111, 112) are gathered on the right side from the viewpoint of the object to be photographed.

[0108] Referring to step S520 of FIG. 5, the electronic device (100) obtains information on the positional relationship of the plurality of cameras based on the recognized shooting direction and the arrangement information of the plurality of cameras. The arrangement information of the plurality of cameras regarding the direction and position of the first camera (111) and the second camera (112) on the electronic device (100) may be stored in advance. The processor (130) may obtain information on the positional relationship of the plurality of cameras (111, 112) based on the recognized shooting direction and the arrangement information of the plurality of cameras (111, 112). Referring also to the embodiment illustrated in FIG. 6A, among the plurality of cameras (111, 112), the first camera (111) may be arranged at the upper portion of the electronic device (100), and the second camera (112) may be arranged below the first camera (111). The processor (130) can obtain position information that, based on the recognized shooting direction and the arrangement information of the plurality of cameras (111, 112), in a shooting direction in which the electronic device (100) is rotated 90° in the left direction, the plurality of cameras (111, 112) are concentrated in the right direction, and the first camera (111) is located relatively further to the right than the second camera (112).

[0109] Referring to step S530 of FIG. 5, the electronic device (100) identifies a camera that is positioned farthest from the center line of the electronic device (100) among the plurality of cameras based on information regarding the positional relationship of the plurality of cameras. Referring also to the embodiment illustrated in FIG. 6A, among the plurality of cameras (111, 112), the camera that is positioned farthest to the right with respect to the virtual center line (l) is the first camera (111). In this case, the processor (130) can identify the first camera (111) that is positioned farthest from the center line (l) among the plurality of cameras (111, 112) based on information regarding the positional relationship of the plurality of cameras (111, 112).

[0110] Referring to step S540 of FIG. 5, the electronic device (100) selects an image captured by an identified camera among the left image and the right image as a reference image. Referring also to operation ② of FIG. 6a, the electronic device (100) captures an object using the first camera (111) in a shooting direction rotated 90° to the left, thereby capturing a right image (i R ) is obtained, and the left image (i) is captured by taking a picture of the object using the second camera (112). L ) can be obtained. The processor (130) of the electronic device (100) can obtain the left image (i L ) and right image (i R ) The right image (i) captured by the first camera (111) identified as the camera most distant from the center line (l) R ) as the reference image for view synthesis (i ref ) can be selected as a source camera. In one embodiment of the present disclosure, the processor (130) identifies the second camera (112) positioned adjacent to the center line (l) as a source camera, and captures the left image (i) captured by the second camera (112) identified as the source camera. L ) can be determined as the target image for synthesis.

[0111] Referring to operation ③ of FIG. 6A, the processor (130) of the electronic device (100) may perform view synthesis on a synthesis target image to obtain a synthesized image (isynthesized). In one embodiment of the present disclosure, the processor (130) may obtain a reference image (i ref ) and the objects included in the synthesized image (isynthesized) can be synthesized by shifting the objects of the synthesized target image to the right so that the objects have a disparity corresponding to the second baseline, which is the distance between the binocular displays of an external device (e.g., a head-mounted display device).

[0112] FIG. 6b shows an electronic device (100) of the present disclosure that captures a left image (i) among stereo images based on the shooting direction and arrangement information of multiple cameras (111, 112). L ) as a reference image (i ref ) and select the right image (i R ) is a drawing illustrating an embodiment of determining a target image for synthesis.

[0113] FIG. 6b is a photographing direction in which the user's grip posture of the electronic device (100) is opposite to that of the embodiment illustrated in FIG. 6a, and a reference image (i) is captured according to the photographing direction. ref ) and the target image for synthesis is determined in reverse, so the description is omitted.

[0114] Referring to operation ① of FIG. 6b, when the electronic device (100) is rotated 90° to the right by the user and held, the processor (130, see FIG. 4) of the electronic device (100) recognizes the shooting direction based on the measurement value by the gravity sensor, and can identify that multiple cameras (111, 112) are gathered on the left side from the viewpoint of the object to be shot.

[0115] The processor (130) can obtain information about the positional relationship of the plurality of cameras (111, 112) based on the recognized shooting direction and the arrangement information of the plurality of cameras (111, 112) on the electronic device (100). In the embodiment illustrated in FIG. 6B, the processor (130) can obtain positional information that, based on the recognized shooting direction and the arrangement information of the plurality of cameras (111, 112), in a shooting direction in which the electronic device (100) is rotated 90° in the right direction, the plurality of cameras (111, 112) are concentrated in the left direction, and the first camera (111) is located relatively further to the left than the second camera (112).

[0116] Referring to operation ② of FIG. 6b, the electronic device (100) captures an object using the first camera (111) to create a left image (i L ) is obtained, and the right image (i) is captured by taking a picture of the object using the second camera (112). R ) is obtained, and the left image (i) is obtained based on information about the positional relationship of multiple cameras (111, 112). L ) and right image (i R ) among the reference images (i) ref ) can be selected. In one embodiment of the present disclosure, the processor (130) identifies the camera positioned furthest from the center line (l) of the electronic device (100) and selects the left image (i L ) and right image (i R ) is the image acquired by the identified camera as the reference image (i ref ) can be selected. In the embodiment illustrated in Fig. 6b, among the plurality of cameras (111, 112), the camera that is arranged farthest to the right from the virtual center line (l) is the first camera (111). In this case, the processor (130) identifies the first camera (111) that is arranged farthest from the center line (l) among the plurality of cameras (111, 112) based on the information about the positional relationship of the plurality of cameras (111, 112), and the left image (i) acquired by the first camera (111) L ) as a reference image (i ref ) can be selected. In one embodiment of the present disclosure, the processor (130) identifies the second camera (112) positioned adjacent to the center line (l) as a source camera, and captures the right image (i) captured by the second camera (112) identified as the source camera. R ) can be determined as the target image for synthesis.

[0117] Referring to operation ③ of FIG. 6b, the processor (130) of the electronic device (100) may perform view synthesis on a synthesis target image to obtain a synthesized image (isynthesized). In one embodiment of the present disclosure, the processor (130) may obtain a reference image (i ref ) and the objects included in the synthesized image (isynthesized) can be synthesized by shifting the objects of the synthesized target image to the left so that the objects have a disparity corresponding to the second baseline, which is the distance between the binocular displays of an external device (e.g., a head-mounted display device).

[0118] The electronic device (100) according to the embodiment illustrated in FIGS. 5, 6a, and 6b recognizes the shooting direction according to the user's grip posture based on the measurement value obtained using the gravity sensor, and generates stereo images (i) based on the shooting direction and the arrangement information of the plurality of cameras (111, 112). L , i R ) among the reference images (i) ref ) and determine a synthesis target image, and perform view synthesis on the synthesis target image to obtain a synthesis image (isynthesized), thereby minimizing the occurrence of artifacts due to view synthesis and providing a technical effect of improving the quality of the image. In addition, since the user (photographer) is likely to look at a real scene in the central area of ​​the electronic device (100) and photograph an object, the electronic device (100) according to one embodiment of the present disclosure is in line with the intention of the user (photographer) and can provide an effect that a person viewing the image through an external device (e.g., a head-mounted display device) views the object from the same position as the photographer, thereby improving the immersion in viewing the image.

[0119] FIG. 7 is a flowchart illustrating a method for an electronic device (100) according to one embodiment of the present disclosure to select a reference image based on location information of a main object in an image.

[0120] Steps S710 to S730 illustrated in FIG. 7 are steps that specify the operation of step S220 of FIG. 2. Step S710 illustrated in FIG. 7 may be performed after the operation of step S210 of FIG. 2 is performed. After the operation according to step S730 illustrated in FIG. 7 is performed, the operation of step S230 of FIG. 2 may be performed.

[0121] In step S710, the electronic device (100) performs image segmentation to recognize at least one object from at least one of the first image and the second image. In the present disclosure, 'image segmentation' refers to an image processing technique that classifies the class or category of an object in an original image, distinguishes the object from other objects or a background image in the image based on the classification result, and segments the object based on the distinguished outline (outlier). In one embodiment of the present disclosure, the processor (130, see FIG. 4) of the electronic device (100) can segment at least one object from an image using a deep neural network model trained to classify a plurality of objects into labels, classes, or categories. Since image segmentation is a well-known technique to those skilled in the art, a detailed description thereof will be omitted.

[0122] FIG. 8 is a diagram showing an electronic device (100) according to an embodiment of the present disclosure, which is an original image (i orig ), segmentation image (i seg ), and depth map image (i depth) is a drawing illustrating an operation of recognizing a main object (800) from an image using a deep neural network model. Referring to step S710 of FIG. 7 together with the embodiment illustrated in FIG. 8, the processor (130) of the electronic device (100) recognizes an original image (i) using a deep neural network model. orig ) to perform image segmentation on the segmented image (i seg ) can be obtained. Segmentation image (i seg ) is distinguished from other objects or background images through outlines and includes a plurality of segmented objects. The processor (130) segments the image (i seg ) can recognize multiple objects.

[0123] Referring back to FIG. 7, in step S720, the electronic device (100) determines a main object based on at least one of the position, direction, size, depth, and movement of at least one recognized object. In one embodiment of the present disclosure, the electronic device (100) may determine, among the at least one recognized object, an object that is closest to the camera, is in a specific direction (e.g., to the right, left, up, down, etc.), is large in size, or is dynamic and moves the most, as a key object, i.e., a main object.

[0124] In one embodiment of the present disclosure, the processor (130) of the electronic device (100) can recognize at least one object based on the segmentation image and depth information of the object, and determine a main object among the at least one recognized object. Referring to the embodiment illustrated in FIG. 8, the processor (130) obtains depth information of the object and generates a depth map image (i depth) can be obtained. For example, the processor (130) measures the depth value of an object based on the disparity between pixels included in each stereo image acquired using a plurality of cameras (111, 112, see FIG. 4) through a stereo imaging method to obtain a depth map image (i depth ) can be obtained. However, it is not limited thereto, and the processor (130) uses SLAM (Simultaneous Localization and Mapping) technology to obtain a three-dimensional depth map image (i) of objects included in the surrounding space. depth ) can be obtained. In addition, for example, the electronic device (100) includes a depth sensor composed of a ToF (Time-of-Flight) sensor or a LiDAR (Light wave Detection and Ranging) sensor, and measures the depth values ​​of objects in the surrounding space using the depth sensor to obtain a depth map image (i depth ) may be obtained. The processor (130) may obtain a segmentation image (i seg ) to recognize multiple objects and a depth map image (i depth ) to obtain depth values ​​for multiple objects, and calculate the average value of the obtained depth values, so that the object with the smallest average value, i.e., the smallest depth value, among the multiple objects can be determined as the main object (800).

[0125] Referring back to FIG. 7, in step S730, the electronic device (100) selects a reference image among the first image and the second image based on the location information at which the main object is arranged within the image. In one embodiment of the present disclosure, if the main object is positioned with a bias toward the first direction with respect to the center of the image, the electronic device (100) may select an image in which the main object is positioned closest to the center of the image among the first image and the second image as the reference image. That is, the electronic device (100) may determine an image in which the main object is positioned most toward the first direction among the first image and the second image as the target image for synthesis. If the main object is positioned with a bias toward the first direction, for example, to the right, the electronic device (100) may select an image in which the main object is positioned relatively in the center region among the first image and the second image as the reference image, and may determine an image in which the main object is positioned relatively to the right as the target image for synthesis. In the opposite case, that is, when the main object is located to the left, the electronic device (100) may select an image in which the main object is located in a relatively central area among the first image and the second image as a reference image, and determine an image in which the main object is located to the relatively left as a target image for synthesis.

[0126] FIG. 9A is a diagram illustrating an embodiment in which an electronic device (100) of the present disclosure determines a synthesis target image based on location information of a main object (900L, 900R) recognized in an image and performs view synthesis on the synthesis target image. Referring to step S730 of FIG. 7 together with the embodiment illustrated in FIG. 9A, the left image (i L ) and the right image (i R ) constitutes a stereo image, and the left image (i L ) and right image (i R) from which the main objects (900L, 900R) can be recognized. The processor (130, see Fig. 4) of the electronic device (100) generates a stereo image (i L , i R ) can recognize the location of the main object (900L, 900R) recognized from the image, and identify that the main object (900L, 900R) is positioned on the right side based on the center area of ​​the image. In this case, the processor (130) can identify that the left image (i L ) and the right image (i R ) is the right image (i) in which the main objects (900L, 900R) are relatively more to the right. R ) and identify the identified right image (i R ) can be determined as a synthesis target image. The processor (130) determines the left image (i L ) and the right image (i R ) is the left image (i) in which the main objects (900L, 900R) are relatively located in the central area of ​​the image. L ) and identify the identified left image (i L ) as a reference image (i ref ) can be determined.

[0127] The processor (130) is a reference image (i ref ) and the object (900R') included in the synthesized image (isynthesized) can be synthesized by shifting the object (900R) in the synthesized target image to the left so that the object (900L) included in the synthesized image and the object (900R') included in the synthesized image have a disparity corresponding to the second baseline, which is the distance between the binocular displays of an external device (e.g., a head-mounted display device), thereby obtaining the synthesized image (isynthesized).

[0128] FIG. 9B is a diagram illustrating an embodiment in which an electronic device (100) of the present disclosure determines a synthesis target image based on location information of a main object (900L, 900R) recognized in an image and performs view synthesis on the synthesis target image. Referring to step S730 of FIG. 7 together with the embodiment illustrated in FIG. 9B, the left image (i L ) and the right image (i R ) constitutes a stereo image, and the left image (i L ) and right image (i R ) from which the main objects (900L, 900R) can be recognized. The processor (130, see Fig. 4) of the electronic device (100) generates a stereo image (i L , i R ) can recognize the location of the main object (900L, 900R) recognized from the image, and identify that the main object (900L, 900R) is positioned on the left side based on the center area of ​​the image. In this case, the processor (130) can identify that the left image (i L ) and the right image (i R ) is the left image (i) in which the main objects (900L, 900R) are relatively more to the left. L ) and identify the identified left image (i L ) can be determined as a synthesis target image. The processor (130) determines the left image (i L ) and the right image (i R ) is the right image (i) in which the main objects (900L, 900R) are relatively located in the central area of ​​the image. R ) and identify the identified right image (i R ) as a reference image (i ref ) can be determined.

[0129] The processor (130) is a reference image (i ref) and the object (900L') included in the synthesized image (isynthesized) can be synthesized by shifting the object (900L) in the synthesized target image to the left so that the object (900R) included in the synthesized image and the object (900L') included in the synthesized image have a disparity corresponding to the second baseline, which is the distance between the binocular displays of an external device (e.g., a head-mounted display device), thereby obtaining the synthesized image (isynthesized).

[0130] In general, viewing convenience can be improved when a main object (e.g., a person, an animal, etc.) in an image is shown in a central area within the image. The electronic device (100) according to the embodiment shown in FIGS. 7, 8, 9a, and 9b recognizes a main object from at least one image among stereo images, and generates a reference image (i) based on the position of the recognized main object. ref ) and determine a synthesis target image, and perform view synthesis to shift a main object within the determined synthesis target image to the center area of ​​the image to obtain a synthesis image (isynthesized), thereby minimizing the occurrence of artifacts due to view synthesis and providing a technical effect of improving user viewing convenience.

[0131] Figure 10 shows the original stereo image (i L , i R ) and stereo image (i) after view synthesis is performed with disparity (d1) in ref , is a diagram showing the parallax (d2) in the synthesized form.

[0132] Referring to Figure 10, the original stereo image (i L , i R ) may have a first parallax (d1). For example, the left image (i) of the original stereo image L) is aligned with a specific pixel of the object (1000L) contained in the right image (i R ) can be separated by the first disparity (d1). Since the disparity in stereo images is inversely proportional to the depth value of the object, the left image (i L ) and the object (1000L) included in the right image (i R ) may have a depth value corresponding to the reciprocal of the first parallax (d1).

[0133] The processor (130, see FIG. 4) of the electronic device (100) can obtain a synthesized image (isynthesized) by performing view synthesis on the synthesis target image so that the parallax corresponds to the baseline, which is the distance between the binocular displays of an external device (e.g., a head-mounted display device). Since the baseline between the binocular displays of the external device is larger than the baseline between the multiple cameras of the electronic device (100), the processor (130) can perform view synthesis to increase the parallax of an object with a small depth value, i.e., an object positioned relatively close to the camera. The stereo image (i) on which view synthesis of FIG. 10 is performed ref , issynthesized), the processor (130) generates a reference image (i ref ) is not synthesized, and the right image (i) is determined as the synthesis target image. R ) can be obtained by performing view synthesis to shift the object (1000R) in the left direction. As the object (1000R) is moved by view synthesis, the object (1000R') included in the synthesized image (isynthesized) is shifted to the reference image (i ref) is separated from the object (1000L) by a second parallax (d2), and an artifact (1000a) may be generated. Specifically, a reference image (i) that is aligned with a specific pixel of an object (1000R') included in the synthesized image (isynthesized) ref ) may be larger than the first disparity (d1) before view synthesis, as the distance between pixels of the object (1000L) is the second disparity (d2). Since the disparity in stereo images is inversely proportional to the depth value of the object, the depth value of the object may decrease as the disparity increases from the first disparity (d1) to the second disparity (d2) after view synthesis. In other words, since an object closer to the camera appears closer, the sense of depth and the user's immersion may be further improved.

[0134] In the embodiment illustrated in FIG. 10, a method is described for obtaining a synthesized image (isynthesized) by maintaining the original image state without performing synthesis on the reference image (iref) and performing view synthesis only on the synthesis target image. However, the present disclosure is not limited thereto, and an electronic device (100) according to an embodiment of the present disclosure may obtain a reference image (i) as well as a synthesis target image. ref ) can also perform view synthesis on the reference image (i ref ) and a specific embodiment of performing view synthesis for both the synthesis target image and the synthesis target image will be described in detail in FIGS. 11 and 12.

[0135] FIG. 11 is a diagram illustrating an electronic device (100) according to an embodiment of the present disclosure that generates an original stereo image (i L , i R ) is a drawing showing a stereo image (iL-synthesized, iR-synthesized) according to the result of view synthesis.

[0136] Referring to Figure 11, the original stereo image is the left image (i L) and right image (i R ) and the left image (i L ) and the right image (i R ) each contains an object (1100L, 1100R). In one embodiment of the present disclosure, the processor (130, see FIG. 4) of the electronic device (100) processes the left image (i) of the original stereo image. L ) as the reference image, and the right image (i R ) can be determined as a synthesis target image. However, this is for convenience of explanation, and the present disclosure is not limited thereto. The processor (130) determines the right image (i R ) as the reference image, the left image (i L ) can also be determined as the target image for synthesis.

[0137] The processor (130) can move both the object (1100L) included in the reference image and the object (1100R) included in the synthesis target image so that the object (1100L) included in the reference image and the object (1100R) included in the synthesis target image have a disparity corresponding to a baseline, which is a distance between the binocular displays of an external device (e.g., a head-mounted display device). In one embodiment of the present disclosure, the processor (130) can perform view synthesis to move the object (1100L) included in the reference image in a first direction (rightward in the embodiment illustrated in FIG. 11) and move the object (1100R) included in the synthesis target image in a direction opposite to the first direction (leftward in the embodiment illustrated in FIG. 11), thereby obtaining a synthesis image (iL-synthesized, iR-synthesized).

[0138] As view synthesis is performed, the object (1100L') included in the left synthesized image (iL-synthesized) is compared to the original image (left image (i L)) was moved by the first distance compared to the object (1100L), and an artifact (1100a-1) was generated due to the movement of the object (1100L'). The object (1100R') included in the right synthetic image (iR-synthesized) is the original image (right image (i R )) was moved by the second distance compared to the object (1100R), and an artifact (1100a-2) was generated due to the movement of the object (1100R'). As in the example illustrated in Fig. 11, the left image (i L ) is selected as a reference image, the size of the first distance moved by the object (1100L') included in the left synthesized image (iL-synthesized) obtained through view synthesis for the reference image may be smaller than the size of the second distance moved by the object (1100R') included in the right synthesized image (iR-synthesized). That is, the distance moved by the object during view synthesis for the image selected as the reference image may be smaller than the distance moved by the object due to view synthesis for the synthesis target image. In the opposite embodiment, that is, the right image (i R ) is selected as the reference image, and the left image (i L ) is determined as the target image for synthesis, the size of the first distance moved by the object (1100L') included in the left synthesized image (iL-synthesized) may be greater than the size of the second distance moved by the object (1100R') included in the right synthesized image (iR-synthesized).

[0139] FIG. 12 is a conceptual diagram illustrating an operation of an electronic device (1000) according to one embodiment of the present disclosure to perform view synthesis according to a synthesis target image.

[0140] Referring to a in Fig. 12, the original stereo image is the left image (i L ) and right image (i R ) may be included.

[0141] In the embodiment shown in b of FIG. 12, the electronic device (100) selects the left image (i) of the original stereo image L ) as a reference image (i ref ) and select the right image (i R ) is determined as a synthesis target image, and a synthesis image (isynthesized) can be obtained by performing view synthesis on the synthesis target image.

[0142] Referring to the embodiment illustrated in c of FIG. 12, the electronic device (100) selects the right image (i) from the original stereo image. R ) as a reference image (i ref ) and select the left image (i L ) is determined as a synthesis target image, and a synthesis image (isynthesized) can be obtained by performing view synthesis on the synthesis target image.

[0143] Referring to the embodiment illustrated in d of FIG. 12, the electronic device (100) selects the left image (i) of the original stereo image L ) and right image (i R ) can be selected as the composite target image and view synthesis can be performed. In this case, the left image (i) is selected due to view synthesis. L ) are moved to the left, and the object contained in the right image (i R ) can be moved to the right and composited.

[0144] Referring to the embodiment illustrated in FIG. 12e, the electronic device (100) selects the left image (i) of the original stereo image. L ) as the reference image, and select the right image (i R) is determined as a synthesis target image, and a left synthesis image (iL-synthesized) and a right synthesis image (iR-synthesized) can be obtained by performing view synthesis on both the reference image and the synthesis target image. In one embodiment of the present disclosure, the right image (i) determined as a synthesis target image by view synthesis R ) is the first distance moved by the object contained in the left image (i) selected as the reference image. L ) may be longer than the second distance moved.

[0145] Referring to the embodiment shown in Fig. 12f, the electronic device (100) selects the right image (i) from the original stereo image. R ) as the reference image, and select the left image (i L ) is determined as a synthesis target image, and a left synthesis image (iL-synthesized) and a right synthesis image (iR-synthesized) can be obtained by performing view synthesis on both the reference image and the synthesis target image. In one embodiment of the present disclosure, the left image (i) determined as a synthesis target image by view synthesis L ) is the first distance moved by the object included in the right image (i) selected as the reference image. R ) may be longer than the second distance moved.

[0146] FIG. 13 is a block diagram illustrating components of a head-mounted display device (200) according to one embodiment of the present disclosure.

[0147] The head mounted display (HMD) device (200) illustrated in FIG. 13 is a device worn on the user's head and provides a virtual reality experience to the user. Although the head mounted display device (200) is illustrated in FIG. 13, the present disclosure is not limited thereto. In one embodiment of the present disclosure, the head mounted display device (200) may be replaced with an augmented reality helmet (ARH) that provides an AR experience, a face mounted display (FMD) device worn on the user's face, or AR glasses (ARG) in the shape of glasses.

[0148] Referring to FIG. 13, the head mounted display device (200) may include a communication interface (210), a display unit (220), a processor (230), and a memory (240). The communication interface (210), the display unit (220), the processor (230), and the memory (240) may be electrically and / or physically connected to each other, respectively. In FIG. 13, only essential components for explaining the function and / or operation of the head mounted display device (200) are illustrated, and the components included in the head mounted display device (200) are not limited as illustrated in FIG. 13. In one embodiment of the present disclosure, the head mounted display device (200) may further include a battery that supplies driving power to the communication interface (210), the display unit (220), and the processor (230). In one embodiment of the present disclosure, the head mounted display device (200) may further include a camera that photographs an object in a real scene. The camera may include, for example, a left-eye camera positioned adjacent to the left-eye display (220L) and a right-eye camera positioned adjacent to the right-eye display (220R). However, the camera is not limited thereto, and the camera may include three or more cameras.

[0149] The communication interface (210) is a hardware device configured to perform data communication with an external device and / or server. The communication interface (210) may be configured as a device that performs data communication with an external device or server using at least one of data communication methods including, for example, Wireless LAN, Wi-Fi, Wi-Fi Direct, Bluetooth, Bluetooth Low Energy (BLE), infrared Data Association (IrDA), Near Field Communication (NFC), Wireless Broadband Internet (Wibro), World Interoperability for Microwave Access (WiMAX), Shared Wireless Access Protocol (SWAP), Wireless Gigabit Alliance (WiGig), and RF communication. The external device is a device that is connected to the head-mounted display device (200) through a communication network, and may be, for example, a mobile device such as a smart phone or a tablet PC.

[0150] In one embodiment of the present disclosure, the communication interface (210) can connect to a mobile device (e.g., 'electronic device (100, see FIGS. 3 and 4)') under the control of the processor (230) and transmit and receive data. The communication interface (210) can pair with the electronic device (100) via a short-range wireless communication network, such as Bluetooth, BLE, or Wi-Fi Direct, and receive a stereo image including a first image and a second image from the electronic device (100). However, the present disclosure is not limited thereto, and in one embodiment of the present disclosure, the head-mounted display device (200) can also obtain a stereo image including a left-eye image and a right-eye image by photographing an object using a camera.

[0151] The display unit (220) is configured to display a reference image and a composite image. The display unit (220) may include a left-eye display (220L) that displays an image toward the user's left eye and a right-eye display (220R) that displays an image toward the user's right eye. The left-eye display (220L) and the right-eye display (220R) may be arranged to be spaced apart from each other by a second baseline. The 'second baseline' may be an average value of the distance between the user's left and right eyes, for example, 6.5 cm. However, the present invention is not limited thereto.

[0152] The left-eye display (220L) and the right-eye display (220R) may be configured with at least one of a liquid crystal display, a thin film transistor-liquid crystal display, an organic light-emitting diode, a flexible display, a 3D display, and an electrophoretic display. For example, when the head-mounted display device (200) is implemented as an augmented reality glass, the left-eye display (220L) and the right-eye display (220R) may be configured with a lens optical system and may include a waveguide and an optical engine. The optical engine may be configured with a projector that generates light of a virtual object composed of a virtual image and projects the light onto a waveguide. The optical engine may include, for example, an image panel, an illumination optical system, a projection optical system, etc. In one embodiment of the present disclosure, the optical engine can generate light of image data of each of a reference image and a composite image, and display the reference image and the composite image by projecting the light into a waveguide.

[0153] The processor (230) can execute one or more instructions of a program stored in the memory (240). The processor (230) may be composed of hardware components that perform arithmetic, logic, and input / output operations, as well as image processing. Although the processor (230) is illustrated as a single element in FIG. 13 , it is not limited thereto. In one embodiment of the present disclosure, the processor (230) may be composed of one or more elements.

[0154] The processor (230) may include various processing circuits and / or multiple processors. For example, the term "processor" as used herein, including the claims, may include various processing circuits, including at least one processor. One or more processors in at least one processor may be configured to perform various functions described herein, individually and / or collectively, in a distributed fashion. As used herein, "processor," "at least one processor," and "one or more processors" may be configured to perform various functions. However, these terms encompass, without limitation, situations where one processor performs some of the functions and other processor(s) perform other parts of the functions, and situations where a single processor may perform all of the functions. Furthermore, at least one processor may include a combination of processors that perform various functions of the disclosed functions in a distributed manner. At least one processor may execute program instructions to achieve or perform various functions.

[0155] The processor (230) may be implemented as a general-purpose processor such as a CPU (Central Processing Unit), an AP (Application Processor), a DSP (Digital Signal Processor), a graphics-only processor such as a GPU (Graphics Processing Unit), a VPU (Vision Processing Unit), or an AI-only processor such as an NPU (Neural Processing Unit), for example. The processor (230) may be controlled to process input data according to predefined operating rules or an AI model. Alternatively, if the processor (230) is an AI-only processor, the AI-only processor may be designed with a hardware structure specialized for processing a specific AI model.

[0156] The memory (240) may be configured as at least one type of storage medium, for example, a flash memory type, a hard disk type, a multimedia card micro type, a card type memory (e.g., SD or XD memory, etc.), a RAM (Random Access Memory), a SRAM (Static Random Access Memory), a ROM (Read-Only Memory), an EEPROM (Electrically Erasable Programmable Read-Only Memory), a PROM (Programmable Read-Only Memory), or an optical disk.

[0157] The memory (240) may store instructions related to functions and / or operations for the head mounted display device (200) to select a reference image among stereo images, determine a synthesis target image, and perform view synthesis for the synthesis target image. In one embodiment of the present disclosure, the memory (240) may store at least one of instructions, an algorithm, a data structure, a program code, and an application program that can be read by the processor (230). The instructions, algorithms, data structures, and program codes stored in the memory (240) may be implemented in a programming or scripting language such as, for example, C, C++, Java, or an assembler.

[0158] The memory (240) may store instructions, algorithms, data structures, or program codes related to the reference image selection module (242) and the view synthesis module (244). A 'module' included in the memory (240) refers to a unit that processes a function or operation performed by the processor (230), and this may be implemented as software such as instructions, algorithms, data structures, or program codes. In one embodiment of the present disclosure, the memory (240) may also include image data storage (246).

[0159] The processor (230) may be implemented by executing instructions or program codes stored in the memory (240). Hereinafter, the functions and / or operations performed by the processor (230) by executing instructions or program codes of each of the plurality of modules stored in the memory (240) will be described in detail.

[0160] The processor (230) may receive stereo images including a first image and a second image from an external device, for example, an electronic device (100, see FIGS. 3 and 4), and determine the received first image and the second image as a left-eye image displayed through the left-eye display (220L), and the second image as a right-eye image displayed through the right-eye display (220R). However, the present invention is not limited thereto, and the processor (230) may determine the first image as a right-eye image displayed through the right-eye display (220R), and may determine the second image as a left-eye image displayed through the left-eye display (220L).

[0161] The reference image selection module (242) is configured with commands or program codes for executing a function and / or operation of selecting a reference image for view synthesis among a plurality of images constituting a stereo image. The function and / or operation of the reference image selection module (242) illustrated in FIG. 13 is the same as that of the reference image selection module (142, see FIG. 4) illustrated in FIG. 4, and therefore, a description overlapping with the description of the reference image selection module (142) is omitted. The processor (230) can select any one image among the received stereo images as a reference image by executing the commands or program codes of the reference image selection module (242).

[0162] The processor (230) may select a reference image among stereo images based on shooting direction information regarding whether the electronic device (100) is held vertically or horizontally when shooting to obtain a stereo image, and arrangement information of a plurality of cameras (111, 112, see FIGS. 3 and 4) included in the electronic device (100). In one embodiment of the present disclosure, the processor (230) may receive information regarding the shooting direction of the electronic device (100) from the electronic device (100) through the communication interface (210). In one embodiment of the present disclosure, the processor (230) may receive arrangement information regarding the direction and position in which the plurality of cameras (111, 112) are arranged on the electronic device (100) through the communication interface (210). However, it is not limited thereto, and the memory (240) of the head mounted display device (200) may have arrangement information of a plurality of cameras (111, 112) included in the electronic device (100) stored in advance.

[0163] The processor (230) can recognize the shooting direction of the electronic device (100) based on information about the shooting direction received from the electronic device (100). For example, when the electronic device (100) is held horizontally by the user and a picture is taken, the processor (230) can recognize the shooting direction based on the information about the shooting direction and the arrangement information of the plurality of cameras (111, 112), whether the picture is taken in a direction in which the plurality of cameras (111, 112) included in the electronic device (100) are positioned to the left or in a direction in which the plurality of cameras (111, 112) are positioned to the right. The processor (230) can select a reference image to serve as a basis for view synthesis among the left-eye image and the right-eye image based on the recognized shooting direction and the arrangement information of the plurality of cameras (111, 112) included in the electronic device (100). The processor (230) can determine the remaining images that are not selected as reference images among the left-eye image and the right-eye image as the synthesis target images.

[0164] The processor (230) may recognize a main object from at least one of the stereo images, and select a reference image based on position information of the main object within the image. In one embodiment of the present disclosure, the processor (230) may perform image segmentation to recognize at least one object from at least one of the left-eye image and the right-eye image, and determine the main object based on at least one of the position, direction, size, depth, and movement of the recognized at least one object. The processor (230) may determine a reference image and a synthesis target image based on position information regarding where the main object is located within the left-eye image and the right-eye image. In one embodiment of the present disclosure, when the main object is located in a first direction with respect to the center of the image, the processor (230) may select an image in which the main object is located closer to the center of the image among the left-eye image and the right-eye image as a reference image, and determine the remaining images among the left-eye image and the right-eye image that are not selected as reference images as synthesis target images. For example, if the main object is positioned to the left within the image, the processor (130) may select an image in which the main object is positioned closest to the center of the image among the left-eye image and the right-eye image as a reference image, and determine the left-eye image as the target image for synthesis. For the opposite example, if the main object is positioned to the right within the image, the processor (130) may determine the right-eye image as the target image for synthesis.

[0165] The view synthesis module (244) is configured with commands or program codes for executing functions and / or operations to obtain a synthesized image by performing view synthesis on a synthesis target image based on depth information of an object. The functions and / or operations of the view synthesis module (244) are the same as those of the view synthesis module (144, see FIG. 4) illustrated in FIG. 4, and therefore, any description overlapping with the description of the view synthesis module (144) will be omitted. The processor (230) can perform view synthesis on a synthesis target image by executing commands or program codes of the view synthesis module (244), thereby obtaining a synthesized image.

[0166] The plurality of cameras (111, 112) included in the electronic device (100) are arranged to be spaced apart from each other by a first baseline, and the first image (left-eye image) acquired through the first camera (111) of the electronic device (100) and the second image (right-eye image) acquired through the second camera (112) may have a disparity corresponding to the first baseline. The left-eye display (220L) and the right-eye display (220R) of the head-mounted display device (200) are spaced apart by a second baseline (for example, 6.5 cm), and when the left-eye image and the right-eye image are displayed through the left-eye display (220L) and the right-eye display (220R), the user may feel a sense of depth that is different from the sense of depth of an actual object, and this may reduce the user's sense of immersion. The processor (230) can obtain a composite image including an object having a depth value according to parallax corresponding to a second baseline, which is a distance between the left-eye display (220L) and the right-eye display (220R), by performing view synthesis on the composite target image based on depth information of the object. In one embodiment of the present disclosure, the processor (230) can perform view synthesis by shifting an object of the composite target image so that the object included in the composite target image and the object included in the reference image have a depth value according to parallax corresponding to the second baseline, and synthesizing the shifted object with the background. The specific method by which the processor (230) performs view synthesis is the same as the method performed by the processor (140) described with reference to FIGS. 4 to 12, and therefore, redundant descriptions are omitted.

[0167] The processor (230) can display a stereo image including a reference image and a composite image through the display unit (220). In one embodiment of the present disclosure, the processor (230) can display the reference image through the left-eye display (220L) and the composite image through the right-eye display (220R). However, the present disclosure is not limited thereto, and in one embodiment of the present disclosure, the processor (230) can also display the composite image through the left-eye display (220L) and the reference image through the right-eye display (220R).

[0168] The image data storage (246) is a storage device within the memory (240) that stores stereo images received from the electronic device (100) or stereo images previously acquired through a network, such as a website. However, the present invention is not limited thereto, and the image data storage (246) may also store images acquired by photographing an object using a camera included in the head-mounted display device (200).

[0169] Image data storage (246) may be composed of non-volatile memory. Non-volatile memory refers to a storage medium that stores and maintains information even when power is not supplied, and can use the stored information again when power is supplied. Non-volatile memory may include, for example, at least one of flash memory, a hard disk, a solid state drive (SSD), a multimedia card micro type, a card type memory (e.g., SD or XD memory), a read only memory (ROM), a magnetic memory, a magnetic disk, and an optical disk.

[0170] Although the image data storage (246) is illustrated in FIG. 13 as a component included within the memory (240), the present disclosure is not limited to the configuration illustrated in the drawing. In one embodiment of the present disclosure, the image data storage (246) may be configured as a database within the head-mounted display device (200), which is a separate component from the memory (240).

[0171] However, the present disclosure is not limited thereto, and in one embodiment of the present disclosure, the image data storage (246) may be implemented as a web storage or cloud server that is accessible via a network and performs a storage function. In this case, the head-mounted display device (200) may communicate with the web storage or cloud server through a communication interface (210) and perform data transmission and reception. The processor (230) may receive stereo images from the web storage or cloud server.

[0172] FIG. 14 is a flowchart illustrating a method for a head-mounted display device (200) according to one embodiment of the present disclosure to perform view synthesis for stereo images.

[0173] In step S1410, the head-mounted display device (200) receives a stereo image including a first image and a second image from an external device. In one embodiment of the present disclosure, the external device may be a mobile device such as a smart phone or a tablet PC. For example, the external device may be the electronic device (100) illustrated in FIGS. 3 and 4. In one embodiment of the present disclosure, the head-mounted display device (200) may pair with the electronic device (100) via a short-range wireless communication network, such as Bluetooth, BLE, or Wi-Fi Direct, and receive a stereo image from the electronic device (100). The head-mounted display device (200) may determine the first image among the received stereo images as a left-eye image displayed through the left-eye display, and the second image as a right-eye image displayed through the right-eye display.

[0174] However, step S1410 is not limited to receiving a stereo image from an external device (e.g., 'electronic device (100, see FIGS. 3 and 4)'). In one embodiment of the present disclosure, the head-mounted display device (200) may further include a camera, and may acquire a stereo image including a left-eye image and a right-eye image by photographing an object using the camera.

[0175] In step S1420, the head mounted display device (200) selects a reference image among the first image and the second image based on at least one of the positional relationship of a plurality of cameras according to the shooting direction of an external device (e.g., 'electronic device (100, see FIGS. 3 and 4)'), the user's dominant eye, or positional information of a main object arranged in an image. A specific method in which the processor (230, see FIG. 13) of the head mounted display device (200) recognizes the shooting direction of the electronic device (100) and selects a reference image among stereo images based on the recognized shooting direction and the arrangement information of the plurality of cameras included in the electronic device (100) is the same as that described in FIG. 13, and therefore, a redundant description thereof will be omitted. The specific method by which the processor (230, see FIG. 13) of the head-mounted display device (200) recognizes a main object from an image and selects a reference image among stereo images based on the location information of the main object is the same as that described in FIG. 13, and therefore, a duplicate description is omitted.

[0176] The head-mounted display device (200) can determine the user's dominant eye by displaying a simulation image through a left-eye display (220L, see FIG. 17) and a right-eye display (220R, see FIG. 17) and receiving a user's input. A specific embodiment in which the head-mounted display device (200) determines the user's dominant eye and determines a reference image and a synthesis target image among stereo images based on the determined dominant eye will be described in detail with reference to FIGS. 16 and 17.

[0177] In step S1430, the head mounted display device (200) determines an image that is not selected as a reference image among the first image and the second image as a synthesis target image.

[0178] In step S1440, the head mounted display device (200) performs view synthesis on the synthesis target image based on the depth information of the object, thereby obtaining a synthesized image including an object having a depth value due to disparity corresponding to a baseline, which is a distance between the left-eye display and the right-eye display. A specific method of obtaining a synthesized image by performing view synthesis on the synthesis target image based on the disparity corresponding to the second baseline between the left-eye display (220L) and the right-eye display (220R), by the processor (230, see FIG. 13) of the head mounted display device (200), is the same as that described in FIG. 13, and therefore, a duplicate description will be omitted.

[0179] In step S1450, the head-mounted display device (200) displays the reference image and the composite image through the left-eye display (220L) and the right-eye display (220R). In one embodiment of the present disclosure, the head-mounted display device (200) may display the reference image on the left-eye display (220L) and the composite image on the right-eye display (220R). However, the present disclosure is not limited thereto, and in one embodiment of the present disclosure, the head-mounted display device (200) may display the reference image on the right-eye display (220R) and the composite image on the left-eye display (220L).

[0180] Unlike the embodiments illustrated in FIGS. 1 to 12, the embodiment illustrated in FIGS. 13 and 14 allows the head-mounted display device (200) to be the subject of an operation of selecting a reference image from among stereo images received from the electronic device (100), determining a synthesis target image, and performing view synthesis on the synthesis target image to obtain a synthesis image. In addition, the head-mounted display device (200) can perform view synthesis on its own and display the obtained synthesis image and the reference image through the display unit (220). Through this, the head-mounted display device (200) according to one embodiment of the present disclosure provides a technical effect of enabling a user to feel the same sense of depth as a real object and enhancing the sense of immersion of a user viewing image content.

[0181] FIG. 15 is a flowchart illustrating an operation method of an electronic device (100) and a head-mounted display device (200) according to one embodiment of the present disclosure.

[0182] Referring to FIG. 15, the head-mounted display device (200) receives information about a shooting direction and position information of a plurality of cameras from the electronic device (100), determines a reference image and a synthesis target image among stereo images (including a first image and a second image) based on the received information about the shooting direction and position information of the plurality of cameras, and performs view synthesis for the synthesis target image. This will be described in detail below.

[0183] The electronic device (100) may be a mobile device such as a smart phone or tablet PC, for example.

[0184] In step S1510, the electronic device (100) acquires a first image and a second image using a plurality of cameras arranged spaced apart from each other by a first baseline. The plurality of cameras included in the electronic device (100) are stereo cameras, and the first image and the second image can constitute a stereo image.

[0185] In step S1520, the electronic device (100) transmits the first image and the second image to the head-mounted display device (200). In one embodiment of the present disclosure, the electronic device (100) may pair with the head-mounted display device (200) via a short-range wireless communication network, such as Bluetooth, BLE, or Wi-Fi Direct, and transmit a stereo image including the first image and the second image to the head-mounted display device (200).

[0186] In step S1530, the electronic device (100) transmits information about the shooting direction when capturing an image and information about the arrangement of a plurality of cameras to the head mounted display device (200). In one embodiment of the present disclosure, the electronic device (100) may transmit shooting direction information, such as whether the electronic device (100) was held vertically or horizontally by the user when capturing an image, and if held horizontally, whether the plurality of cameras were held in a direction in which the plurality of cameras were clustered in the left direction or in a direction in which the plurality of cameras were clustered in the right direction, to the head mounted display device (200). In this case, the electronic device (100) may transmit information about the arrangement of the plurality of cameras to the head mounted display device (200) together with the shooting direction information. However, the present invention is not limited thereto, and in one embodiment of the present disclosure, the head-mounted display device (200) may obtain arrangement information of a plurality of cameras included in the electronic device (100) in advance, and the electronic device (100) may not transmit arrangement information of the plurality of cameras to the head-mounted display device (200).

[0187] In step S1540, the head mounted display device (200) selects a reference image among the first image and the second image based on shooting direction information and arrangement information of multiple cameras.

[0188] In step S1550, the head mounted display device (200) determines the remaining images that are not selected as reference images among the first image and the second image as the synthesis target images.

[0189] In step S1560, the head mounted display device (200) performs view synthesis on a synthesis target image based on depth information of the object, thereby obtaining a synthesis image.

[0190] In step S1570, the head mounted display device (200) displays the reference image and the composite image through the left eye display and the right eye display.

[0191] The specific method of the head mounted display device (200) illustrated in steps S1540 to S1570 is the same as the method described in FIGS. 13 and 14, so redundant description is omitted.

[0192] FIG. 16 is a flowchart illustrating a method for selecting a reference image for view synthesis based on a user's dominant eye by a head-mounted display device (200) according to one embodiment of the present disclosure.

[0193] Steps S1610 to S1640 illustrated in FIG. 16 are steps that embody the operation of step S1420 of FIG. 14. Step S1610 illustrated in FIG. 16 may be performed after the operation of step S1410 of FIG. 14 is performed.

[0194] In step S1610, the head mounted display device (200) alternately displays the original image and the blurred image on the left eye display and the right eye display.

[0195] FIG. 17 is a diagram illustrating an operation of a head-mounted display device (200) according to an embodiment of the present disclosure to determine a user's dominant eye by displaying a simulation image. Referring to step S1610 of FIG. 16 together with the embodiment illustrated in FIG. 17, the head-mounted display device (200) may display a virtual simulation image implemented by software to determine the user's dominant eye. The 'simulated image' may include an original image and a blurred image. In an embodiment of the present disclosure, the head-mounted display device (200) displays the same image through the left-eye display (220L) and the right-eye display (220R), but the original image (i) is displayed on one of the display units of the left-eye display (220L) and the right-eye display (220R). orig ) is displayed on one display section, and a blurred image (i) is displayed on the other display section. blur ) can be displayed.

[0196] In the simulation image, the original image and the blurred image can be displayed alternately over time through the left-eye display (220L) and the right-eye display (220R). For example, the head-mounted display device (200) can display the original image (i) through the left-eye display (220L) at a first time point (t1). orig ) and display a blur image (i) through the right-eye display (220R). blur ) can be displayed. At the second time point (t2), the head mounted display device (200) displays a blur image (i) through the left eye display (220L). blur ) and display the original image (i) through the right-hand display (220R). orig) can be displayed. At the third time point (t3), the head mounted display device (200) can display the original image (i) through the left eye display (220L). orig ) and display a blur image (i) through the right-eye display (220R). blur ) can be displayed.

[0197] Referring back to FIG. 16, in step S1620, the head-mounted display device (200) receives a user input for selecting a display on which the image is clearly visible among the left-eye display and the right-eye display. Referring also to the embodiment illustrated in FIG. 17, the head-mounted display device (200) displays the original image (i) through the left-eye display (220L) and the right-eye display (220R) over time. orig ) and blurred image (i blur ) can be alternately displayed and a query message can be output asking the user which display shows the image more clearly. The head mounted display device (200) can receive the user's response to the query message through the user's hand gesture, touch input, or voice input. However, the present invention is not limited thereto, and the head mounted display device (200) can receive the user's response through any known method.

[0198] Referring to step S1630 of FIG. 16, the head-mounted display device (200) determines the user's dominant eye based on a display selected by a user input. In the present disclosure, the 'dominant eye' refers to the eye that is primarily relied on when receiving visual information among a person's two eyes. Conversely, the eye that is not primarily responsible for visual information is referred to as the secondary eye. In one embodiment of the present disclosure, the head-mounted display device (200) may determine one of the displays based on the user's response to which display shows a clearer image among the left-eye display (220L, see FIG. 17) and the right-eye display (220R, see FIG. 17). The head-mounted display device (200) may determine the user's dominant eye based on the selected display.

[0199] In step S1640, the head-mounted display device (200) selects a reference image based on the dominant eye. For example, if the user's dominant eye is the left eye, the head-mounted display device (200) may select the left-eye image displayed through the left-eye display (220L) as the reference image. For the opposite example, if the user's dominant eye is the right eye, the head-mounted display device (200) may select the right-eye image displayed through the right-eye display (220R) as the reference image.

[0200] In general, when an image viewed through the dominant eye of the user's binocular vision is clear, the image viewed through the binocular vision is perceived as having higher clarity. The head-mounted display device (200) according to the embodiment illustrated in FIGS. 16 and 17 determines the dominant eye of the user through a simulation image, selects a reference image based on the dominant eye, determines the remaining images among the binocular images that are not selected as reference images as synthesis target images, and performs view synthesis on the determined synthesis target images, so that the synthesis image including artifacts is used for the purpose of expressing a sense of depth, thereby providing a technical effect that allows the user to feel the entire image viewed through the binocular vision as clear.

[0201] The present disclosure provides an electronic device (100) that performs view synthesis. According to an embodiment of the present disclosure, the electronic device (100) may include a plurality of cameras (111, 112) arranged spaced apart from each other by a first baseline, a communication interface (150) that performs data communication with an external device, a processor (130) including processing circuitry, and a memory (140) that stores one or more instructions. The one or more instructions are individually or collectively executed by at least one processor (130), so that the electronic device (100) can capture an object using the plurality of cameras (111, 112) to obtain a first image and a second image. By individually or collectively executing the one or more commands by at least one processor (130), the electronic device (100) can select a reference image among the first image and the second image based on the positional relationship of the plurality of cameras (111, 112) according to the shooting direction of the user when shooting an object or the positional information of the main object placed in the image. By individually or collectively executing the one or more commands by at least one processor (130), the electronic device (100) can determine the remaining images that are not selected as reference images among the first image and the second image as a synthesis target image.By individually or collectively executing the one or more commands by at least one processor (130), the electronic device (100) can perform view synthesis on a synthesis target image based on depth information of the object, thereby obtaining a synthesized image including an object having a depth value by disparity corresponding to a second baseline, which is a distance between the left-eye display (220L) and the right-eye display (220R) of the external device (200). By individually or collectively executing the one or more commands by at least one processor (130), the electronic device (100) can control the communication interface (150) to transmit the reference image and the synthesized image to the external device (200) so that the reference image and the synthesized image are displayed by the external device (200).

[0202] In one embodiment of the present disclosure, the electronic device (100) may further include a gravity sensor configured to detect gravity and obtain a measurement value regarding the direction of gravity. By individually or collectively executing the one or more commands by at least one processor (130), the electronic device (100) may recognize a shooting direction based on a posture in which the electronic device is held by a user during shooting based on the measurement value obtained by the gravity sensor. By individually or collectively executing the one or more commands by at least one processor (130), the electronic device (100) may obtain information regarding the positional relationship of the plurality of cameras (111, 112) based on the recognized shooting direction and arrangement information of the plurality of cameras (111, 112) on the electronic device (100). By individually or collectively executing one or more of the above commands by at least one processor (130), the electronic device (100) can select a reference image among the first image and the second image based on information about the positional relationship of the plurality of cameras (111, 112).

[0203] In one embodiment of the present disclosure, the one or more commands are individually or collectively executed by at least one processor (130), so that the electronic device (100) can identify a camera positioned farthest from a center line of the electronic device (100) among the plurality of cameras (111, 112) based on information about the positional relationship of the plurality of cameras (111, 112), and select an image captured by the identified camera among the first image and the second image as a reference image.

[0204] In one embodiment of the present disclosure, the one or more commands are individually or collectively executed by at least one processor (130), so that the electronic device (100) can perform view synthesis by shifting and synthesizing an object of a synthesis target image in a second direction opposite to the first direction so that an object included in a reference image and an object included in a synthesis image have a parallax corresponding to a second baseline when a plurality of cameras (111, 112) are positioned in a first direction with respect to a center line of the electronic device (100).

[0205] In one embodiment of the present disclosure, the one or more commands are individually or collectively executed by at least one processor (130), so that the electronic device (100) can perform image segmentation to recognize at least one object from at least one of a first image and a second image, and determine a main object based on at least one of a position, a direction, a size, a depth, and a movement of the recognized at least one object. The one or more commands are individually or collectively executed by at least one processor (130), so that the electronic device (100) can select a reference image from among the first image and the second image based on position information at which the main object is arranged within the image.

[0206] In one embodiment of the present disclosure, the one or more commands are individually or collectively executed by at least one processor (130), so that the electronic device (100) can select, as a reference image, an image in which the main object is positioned closest to the center of the image among the first image and the second image when the main object is positioned in a first direction relative to the center of the image.

[0207] In one embodiment of the present disclosure, the one or more commands are individually or collectively executed by at least one processor (130), so that the electronic device (100) can obtain a composite image by shifting and synthesizing an object in a composite target image in a second direction opposite to the first direction so that an object included in a reference image and an object included in a composite image have a parallax corresponding to a second baseline.

[0208] In one embodiment of the present disclosure, the external device (200) is a head-mounted display device, and the reference image and the composite image can be displayed by the left-eye display (220L) and the right-eye display (220R) of the head-mounted display device, respectively.

[0209] In one embodiment of the present disclosure, the one or more commands are individually or collectively executed by at least one processor (130), so that the electronic device (100) can obtain a first synthesized image by performing view synthesis that shifts an object included in a synthesized target image in a first direction so that the object has a depth value according to parallax corresponding to a second baseline, and can obtain a second synthesized image by performing view synthesis that shifts an object included in a reference image in a second direction opposite to the first direction.

[0210] In one embodiment of the present disclosure, a first distance by which an object of a synthetic target image moves along a first direction may be greater than a second distance by which an object of a reference image moves along a second direction.

[0211] The present disclosure provides a method for an electronic device (100) to perform view synthesis. The operating method of the electronic device (100) according to one embodiment of the present disclosure may include a step (S210) of acquiring a stereo image including a first image and a second image by photographing an object using a plurality of cameras (111, 112) arranged spaced apart from each other by a first baseline. The operating method of the electronic device (100) according to one embodiment of the present disclosure may include a step (S220) of selecting a reference image among the first image and the second image based on a positional relationship of the plurality of cameras (111, 112) according to a shooting direction of a user when photographing an object or positional information of a main object arranged within the image. The operating method of the electronic device (100) according to one embodiment of the present disclosure may include a step (S230) of determining a remaining image among the first image and the second image that is not selected as a reference image as a synthesis target image. The operating method of the electronic device (100) according to one embodiment of the present disclosure may include a step (S240) of obtaining a synthesized image including an object having a depth value due to disparity corresponding to a second baseline, which is a distance between a left-eye display and a right-eye display of an external device (200), by performing view synthesis on a synthesis target image based on depth information of the object. The operating method of the electronic device (100) according to one embodiment of the present disclosure may include a step (S250) of transmitting a reference image and a synthesized image to an external device (200) so that the reference image and the synthesized image are displayed by the external device (200).

[0212] In one embodiment of the present disclosure, the electronic device (100) may further include a gravity sensor configured to detect gravity and obtain a measurement value regarding the direction of gravity. The step of selecting the reference image (S220) may include a step of recognizing a shooting direction based on a posture in which the electronic device (100) is held by a user during shooting based on the measurement value obtained by the gravity sensor (S510), and a step of obtaining information regarding the positional relationship of the plurality of cameras (111, 112) based on the recognized shooting direction and arrangement information of the plurality of cameras (111, 112) on the electronic device (100) (S520). The step of selecting the reference image (S220) may include a step of selecting the reference image from among the first image and the second image based on the information regarding the positional relationship of the plurality of cameras (111, 112).

[0213] In one embodiment of the present disclosure, the step (S220) of selecting the reference image may include a step (S530) of identifying a camera positioned farthest from a center line of the electronic device (100) among the plurality of cameras (111, 112) based on information regarding the positional relationship of the plurality of cameras (111, 112). The step (S220) of selecting the reference image may include a step (S540) of selecting an image captured by the identified camera among the first image and the second image as the reference image.

[0214] In one embodiment of the present disclosure, in the step (S240) of obtaining the composite image, when the plurality of cameras (111, 112) are positioned in a first direction relative to the center line of the electronic device (100), the electronic device (100) may perform view synthesis by shifting and synthesizing objects of the composite target image in a second direction opposite to the first direction so that the objects included in the reference image and the objects included in the composite image have a parallax corresponding to the second baseline.

[0215] In one embodiment of the present disclosure, the step of selecting a reference image (S220) may include a step of recognizing at least one object from at least one of a first image and a second image by performing image segmentation (S710), and a step of determining a main object based on at least one of a position, a direction, a size, a depth, and a movement of the recognized at least one object (S720). The step of selecting a reference image (S220) may include a step of selecting a reference image from among the first image and the second image (S730) based on location information at which the main object is arranged within the image.

[0216] In one embodiment of the present disclosure, in the step (S220) of selecting the reference image, the electronic device (100) may select, as the reference image, an image in which the main object is positioned closest to the center of the image among the first image and the second image when the main object is positioned in a first direction relative to the center of the image.

[0217] In one embodiment of the present disclosure, in the step (S240) of obtaining the composite image, the electronic device (100) can obtain the composite image by shifting and synthesizing the object in the composite target image in a second direction opposite to the first direction so that the object included in the reference image and the object included in the composite image have a parallax corresponding to the second baseline.

[0218] In one embodiment of the present disclosure, in the step (S240) of obtaining the composite image, the electronic device (100) may obtain a first composite image by performing view synthesis that shifts an object included in a composite target image in a first direction so that the object has a depth value according to parallax corresponding to a second baseline, and may obtain a second composite image by performing view synthesis that shifts an object included in a reference image in a second direction opposite to the first direction.

[0219] In one embodiment of the present disclosure, a first distance by which an object of a synthetic target image moves along a first direction may be greater than a second distance by which an object of a reference image moves along a second direction.

[0220] The present disclosure provides a head mounted display (HMD) device (200) that performs view synthesis. The head mounted display device (200) according to one embodiment of the present disclosure may include a communication interface (210) that pairs with an external device and performs data communication with the external device, a left-eye display (220L) and a right-eye display (220R) spaced apart from each other by a baseline, at least one processor (230) including a processing circuit, and a memory (240) that stores one or more instructions. The one or more instructions are individually or collectively executed by the at least one processor (230), whereby the head mounted display device (200) can control the communication interface (210) to receive a stereo image including a first image and a second image acquired by a plurality of cameras (111, 112) included in the external device. By individually or collectively executing one or more of the above commands by at least one processor (230), the head mounted display device (200) can select a reference image among the first image and the second image based on at least one of the positional relationship of the plurality of cameras (111, 112) according to the shooting direction of the external device, the dominant eye of the user, or the positional information of the main object placed in the image. By individually or collectively executing one or more of the above commands by at least one processor (230), the head mounted display device (200) can determine the remaining image that is not selected as the reference image among the first image and the second image as a synthesis target image.By individually or collectively executing the one or more commands by at least one processor (230), the head mounted display device (200) can perform view synthesis on a synthesis target image based on depth information of the object, thereby obtaining a synthesized image including an object having a different depth value at a disparity corresponding to a baseline. By individually or collectively executing the one or more commands by at least one processor (230), the head mounted display device (200) can display the reference image and the synthesized image through the left eye display (220L) and the right eye display (220R).

[0221] The program executed by the electronic device (100) described in the present disclosure may be implemented as hardware components, software components, and / or a combination of hardware components and software components. The program may be executed by any system capable of executing computer-readable instructions.

[0222] Software may include a computer program, code, instructions, or a combination of one or more of these, which may configure a processing device to do a desired thing or may independently or collectively command a processing device to do a desired thing.

[0223] Software may be implemented as a computer program containing instructions stored on a computer-readable storage medium. Examples of computer-readable storage media include magnetic storage media (e.g., read-only memory (ROM), random-access memory (RAM), floppy disks, hard disks, etc.) and optical readable media (e.g., CD-ROMs, DVDs (Digital Versatile Discs)). The computer-readable storage media may be distributed across network-connected computer systems, so that computer-readable code may be stored and executed in a distributed manner. The media may be readable by a computer, stored in a memory, and executed by a processor.

[0224] A computer-readable storage medium may be provided in the form of a non-transitory storage medium. Here, "non-transitory" simply means that the storage medium does not contain signals and is tangible, but does not distinguish between cases where data is stored semi-permanently or temporarily on the storage medium. For example, a "non-transitory storage medium" may include a buffer in which data is temporarily stored.

[0225] Additionally, programs according to the embodiments disclosed herein may be provided as part of a computer program product. The computer program product may be traded as a commodity between sellers and buyers.

[0226] A computer program product may include a software program, a computer-readable storage medium having the software program stored thereon. For example, the computer program product may be available from a manufacturer of an electronic device (100) or an electronic market (e.g., Samsung Galaxy Store). TM) may include a product in the form of a software program (e.g., a downloadable application) that is distributed electronically. For electronic distribution, at least a portion of the software program may be stored in a storage medium or temporarily created. In this case, the storage medium may be a storage medium of a server of a manufacturer of the electronic device (100), a server of an electronic market, or a relay server that temporarily stores the software program.

[0227] The computer program product may include a storage medium of the server or the storage medium of the electronic device (100) in a system comprising an electronic device (100) and / or a server. Alternatively, if there is a third device (e.g., a 'head-mounted display device (200, see FIGS. 3 and 13)') that is communicatively connected to the electronic device (100), the computer program product may include a storage medium of the third device. Alternatively, the computer program product may include a software program itself that is transmitted from the electronic device (100) to the third device, or from the third device to the electronic device.

[0228] In this case, either the electronic device (100) or the third device may execute the computer program product to perform the method according to the disclosed embodiments. Alternatively, at least one of the electronic device (100) and the third device may execute the computer program product to perform the method according to the disclosed embodiments in a distributed manner.

[0229] For example, the electronic device (100) may execute a computer program product stored in a memory (140, see FIG. 4) to control another electronic device (e.g., a 'head-mounted display device (200, see FIGS. 3 and 13)') that is in communication with the electronic device (100) to perform a method according to the disclosed embodiments.

[0230] As another example, a third device (e.g., a 'head mounted display device (200, see FIGS. 3 and 13)') may execute a computer program product to control an electronic device in communication with the third device to perform a method according to the disclosed embodiment.

[0231] When a third device executes a computer program product, the third device may download the computer program product from the electronic device (100) and execute the downloaded computer program product. Alternatively, the third device may execute a computer program product provided in a pre-loaded state to perform the method according to the disclosed embodiments.

[0232] Although the embodiments described above have been described with limited examples and drawings, those skilled in the art will appreciate that various modifications and variations can be made based on the above description. For example, appropriate results can still be achieved even if the described techniques are performed in a different order than described, and / or components such as the described computer system or modules are combined or combined in a different manner than described, or are replaced or substituted with other components or equivalents.

Claims

1. In an electronic device (100) that performs view synthesis, A plurality of cameras (111, 112) arranged spaced apart from each other by a first baseline; A communication interface (150) that performs data communication with an external device (200); At least one processor (130) comprising processing circuitry; and A memory (130) storing one or more instructions; Including, The electronic device (100) is configured such that the one or more instructions are individually or collectively executed by the at least one processor (130). By photographing an object using the above multiple cameras (111, 112), a first image and a second image are obtained, Selecting a reference image among the first image and the second image based on the positional relationship of the plurality of cameras according to the user's shooting direction when shooting the object or the positional information of the main object placed within the image, Among the first image and the second image, the remaining images that are not selected as the reference images are determined as the target images for synthesis, By performing view synthesis on the synthesis target image based on the depth information of the object, a synthesized image including an object having a depth value due to disparity corresponding to a second baseline, which is a distance between a left-eye display (220L) and a right-eye display (220R) of an external device (200), is obtained. An electronic device (100) that controls the communication interface (150) to transmit the reference image and the composite image to the external device (200) so that the reference image and the composite image are displayed by the external device (200).

2. In paragraph 1, A gravity sensor configured to detect gravity and obtain measurements regarding the direction of gravity; Including more, The electronic device (100) is configured such that the one or more of the above instructions are individually or collectively executed by the at least one processor (130). Based on the measurement value acquired by the gravity sensor, the shooting direction is recognized by the posture in which the electronic device is held by the user during shooting, Based on the recognized shooting direction and the arrangement information of the plurality of cameras (111, 112) on the electronic device (100), information about the positional relationship of the plurality of cameras (111, 112) is obtained. An electronic device (100) that selects the reference image among the first image and the second image based on information about the positional relationship of the plurality of cameras (111, 112).

3. In paragraph 2, The electronic device (100) is configured such that the one or more of the above instructions are individually or collectively executed by the at least one processor (130). Based on the information about the positional relationship of the plurality of cameras (111, 112), the camera positioned farthest from the center line of the electronic device (100) among the plurality of cameras (111, 112) is identified, Selecting an image captured by the identified camera among the first image and the second image as the reference image, An electronic device (100) that performs view synthesis by shifting and synthesizing objects of the synthesis target image in a second direction opposite to the first direction so that the objects included in the reference image and the objects included in the synthesis image have a parallax corresponding to the second baseline when the plurality of cameras (111, 112) are positioned in a first direction with respect to the center line of the electronic device (100).

4. In paragraph 1, The electronic device (100) is configured such that the one or more of the above instructions are individually or collectively executed by the at least one processor (130). By performing image segmentation, at least one object is recognized from at least one of the first image and the second image, determining the main object based on at least one of the position, orientation, size, depth, and movement of the at least one recognized object; An electronic device (100) that selects the reference image among the first image and the second image based on position information at which the main object is placed within the image.

5. In paragraph 4, The electronic device (100) is configured such that the one or more of the above instructions are individually or collectively executed by the at least one processor (130). If the main object is positioned in the first direction relative to the center of the image, the image in which the main object is positioned closest to the center of the image among the first image and the second image is selected as the reference image, An electronic device (100) that obtains the composite image by shifting and synthesizing an object in the composite target image in a second direction opposite to the first direction so that the object included in the reference image and the object included in the composite image have a parallax corresponding to the second baseline.

6. In any one of clauses 1 to 5, The above external device (200) is a head-mounted display device, An electronic device (100), wherein the above reference image and the above composite image are displayed by the left eye display (220L) and the right eye display (220R) of the head mounted display device, respectively.

7. In any one of clauses 1 to 6, The electronic device (100) is configured such that the one or more of the above instructions are individually or collectively executed by the at least one processor (130). An electronic device (100) that obtains a first synthesized image by performing view synthesis that shifts an object included in the synthesized target image in a first direction so that the object has a depth value according to parallax corresponding to the second baseline, and obtains a second synthesized image by performing view synthesis that shifts an object included in the reference image in a second direction opposite to the first direction.

8. In paragraph 7, An electronic device (100), wherein a first distance by which an object of the above-mentioned synthetic target image moves along the first direction is greater than a second distance by which an object of the above-mentioned reference image moves along the second direction.

9. In a method for an electronic device (100) to perform stereo view synthesis, A step (S210) of obtaining a stereo image including a first image and a second image by photographing an object using a plurality of cameras (111, 112) arranged spaced apart from each other by a first baseline; A step (S220) of selecting a reference image among the first image and the second image based on the positional relationship of the plurality of cameras (111, 112) according to the user's shooting direction when shooting the object or the positional information of the main object placed within the image; Step (S230) of determining the remaining images among the first image and the second image that are not selected as the reference images as the target images for synthesis; A step (S240) of obtaining a synthesized image including an object having a depth value by disparity corresponding to a second baseline, which is a distance between the left-eye display and the right-eye display of an external device (200), by performing view synthesis on the synthesized target image based on the depth information of the object; and A step (S250) of transmitting the reference image and the composite image to the external device (200) so that the reference image and the composite image are displayed by the external device (200); A method comprising:

10. In paragraph 9, A gravity sensor configured to detect gravity and obtain measurements regarding the direction of gravity; Including more, The step of selecting the above reference image (S220) is A step (S510) of recognizing a shooting direction based on a posture in which the electronic device (100) is held by the user during shooting based on the measurement value acquired by the gravity sensor; Step (S520) of obtaining information on the positional relationship of the plurality of cameras (111, 112) based on the recognized shooting direction and the arrangement information of the plurality of cameras (111, 112) on the electronic device (100); and A step of selecting the reference image among the first image and the second image based on information about the positional relationship of the plurality of cameras (111, 112); A method comprising:

11. In Article 10, The step of selecting the above reference image (S220) is A step (S530) of identifying a camera positioned farthest from the center line of the electronic device (100) among the plurality of cameras (111, 112) based on information about the positional relationship of the plurality of cameras (111, 112); and A step (S540) of selecting an image captured by the identified camera among the first image and the second image as the reference image; Including, The step (S240) of obtaining the above composite image is A method for performing view synthesis by shifting an object of the synthesis target image in a second direction opposite to the first direction so that an object included in the reference image and an object included in the synthesis image have a parallax corresponding to the second baseline when the plurality of cameras (111, 112) are positioned in a first direction with respect to the center line of the electronic device (100).

12. In paragraph 9, The step of selecting the above reference image (S220) is A step (S710) of recognizing at least one object from at least one of the first image and the second image by performing image segmentation; A step (S720) of determining the main object based on at least one of the position, direction, size, depth, and movement of the at least one recognized object; and A step (S730) of selecting the reference image among the first image and the second image based on the location information where the main object is placed within the image; A method comprising:

13. In paragraph 12, The step of selecting the above reference image (S220) is If the main object is positioned in the first direction relative to the center of the image, the image in which the main object is positioned closest to the center of the image among the first image and the second image is selected as the reference image, The step (S240) of obtaining the above composite image is A method for obtaining the composite image by shifting and synthesizing an object in the composite target image in a second direction opposite to the first direction so that the object included in the reference image and the object included in the composite image have a parallax corresponding to the second baseline.

14. In any one of paragraphs 9 to 13, The step (S240) of obtaining the above composite image is A method for obtaining a first synthetic image by performing view synthesis to shift an object included in the synthetic target image in a first direction so that the object has a depth value according to parallax corresponding to the second baseline, and obtaining a second synthetic image by performing view synthesis to shift an object included in the reference image in a second direction opposite to the first direction.

15. In paragraph 14, A method wherein a first distance by which an object of the above-mentioned synthetic target image moves along the first direction is greater than a second distance by which an object of the above-mentioned reference image moves along the second direction.

Citation Information

Patent Citations

  • Information processing device and information processing method

    JP2013121150A

  • Image display system, image display program, image display device, and image display method

    JP7349808B2

  • Apparatus and method for processing information of multi camera

    KR1020180083245A

  • Tunnel entry barrier using indoor fire hydrant and remote control automatic fire extinguishing system

    KR102169077B1

  • The Electronic Device and the Method for Processing Image

    KR102431488B1