Three-dimensional scene repair based on stereo extraction

By acquiring two-dimensional images from different viewpoints, creating depth maps and reconstructing three-dimensional scenes, the problem of generating unreal appearance shapes and colors in the existing technology is solved, and the effect of efficiently generating real three-dimensional scenes is achieved.

CN120451009APending Publication Date: 2025-08-08SNAP INC
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510522392.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2018-09-27
Filing Date
2019-09-04
Publication Date
2025-08-08

AI Technical Summary

Technical Problem

The prior art is prone to unreal appearance shapes and colors when creating three-dimensional scenes from a limited number of two-dimensional images, and requires a lot of processing time and resources.

Method used

By obtaining two-dimensional images from two different viewpoints, creating depth maps and reconstructing 3D scenes, detecting incomplete areas, identifying replacement information and correcting these areas, and finally rendering the 3D scene from multiple viewpoints.

Benefits of technology

It realizes the generation of more realistic three-dimensional scenes in a shorter time, reducing processing time and resource requirements, while avoiding unreal appearance shapes and color effects.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120451009A_ABST
    Figure CN120451009A_ABST
Patent Text Reader

Abstract

Systems and methods for rendering a three-dimensional (3D) scene with improved visual characteristics from a pair of two-dimensional images from different viewpoints. Creating a three-dimensional scene by obtaining a first two-dimensional (2D) image of a scene object acquired from a first viewpoint and a second two-dimensional image of the scene object acquired from a second viewpoint different from the first viewpoint, creating a depth map from the first and second two-dimensional images, creating a three-dimensional scene from the depth map and the first and second two-dimensional images, detecting an area with incomplete image information in the initial three-dimensional scene, reconstructing the detected area in the three-dimensional scene, determining replacement information and correcting the reconstructed area; and rendering the three-dimensional scene with the modified reconstruction region from the plurality of viewpoints.
Need to check novelty before this filing date? Find Prior Art

Description

CROSS-REFERENCE TO RELATED APPLICATIONS

[0001] This application claims priority to U.S. Provisional Application Serial No. 62 / 737,280, filed on September 27, 2018, entitled “3D Scene Inpainting Using Stereo Extraction,” the entire contents of which are incorporated herein by reference. Technical Field

[0002] The present invention relates to rendering three-dimensional scenes, and more particularly, to systems and methods for rendering a two-dimensional scene with improved visual properties from a pair of two-dimensional images having different viewpoints. Background Art

[0003] A computer system uses a rendering program to present three-dimensional (3D) scene objects on a two-dimensional (2D) display. To create a 3D representation of the scene objects, the computer system obtains geometric information about the scene objects from multiple 2D images taken from different viewpoints. The computer system then creates a depth map based on the obtained geometric information, which is used to create and render the 3D scene on the 2D display.

[0004] A depth map is an image that contains information about the distance from the imaging viewpoint that captures the two-dimensional image of the scene object to the surface of the scene object. This depth is sometimes referred to as Z-depth, which refers to the convention that the central axis of the imager's viewpoint is in the direction of the imager's Z axis, rather than in the absolute Z axis of the scene.

[0005] A computer system presents a three-dimensional scene on a two-dimensional display for viewing and manipulation by a user, and the user is able to manipulate scene objects of the three-dimensional scene by changing the viewpoint of the three-dimensional scene. With respect to the viewpoint of the scene objects, the three-dimensional scene does not include accurate information (e.g., color values, depth values, and / or object values, because some faces of the object / scene do not appear in one or more two-dimensional images obtained from the limited number of viewpoints used to create the three-dimensional scene), and the computer system attempts to complete the scene using information from neighboring pixels. When the computer system attempts to complete the scene (e.g., using information from neighboring pixels), the resulting scene often includes unrealistic-looking shapes and / or colors (e.g., a "stretching" effect of colors).

[0006] To minimize unrealistically shaped and / or colored objects, existing techniques often obtain many more than two 2D images from two different viewpoints (e.g., a panorama of the image). Obtaining a panorama of the image from multiple viewpoints increases the likelihood that at least one viewpoint of the scene object includes information for generating a 3D scene from various viewpoints, thereby improving display accuracy. However, this technique requires a relatively large amount of processing time and power compared to creating a 3D scene from only two 2D images. BRIEF DESCRIPTION OF THE DRAWINGS

[0007] The present invention will be best understood from the following detailed description when read in conjunction with the accompanying drawings, in which like components are given like reference numerals. When multiple similar components are present, a single reference numeral may be assigned to multiple similar components, and lowercase letters may be used to indicate specific components. When referring to a collection of components or to a non-specific component or components, lowercase letters may be omitted. This emphasizes that, according to convention, unless otherwise indicated, the various features of the drawings are not drawn to scale. Rather, the dimensions of various features may be expanded or reduced for clarity. Included in the drawings are the following figures:

[0008] Figure 1A is a perspective view of an example of eyewear including an electronic component and a support structure supporting the electronic component;

[0009] Figure 1B yes Figure 1A a top view of an example of an eyeglass, showing the area defined by the eyeglasses for accommodating the head of a user wearing the eyeglasses;

[0010] Figure 2 yes Figure 1A Block diagram of examples of electronic components supported by an example of glasses, and communications with a personal computing device and a recipient.

[0011] Figure 3A is a flowchart of example steps for rendering a 3D scene;

[0012] Figure 3B is used to reconstruct the first and second 2D images into Figure 3A A flowchart of example steps for creating a 3D scene in FIG.

[0013] Figure 3C is used for Figure 3B A flowchart of example steps for detecting regions of a 3D scene with incomplete information; and

[0014] Figure 3D Is used to determine Figure 3B Flowchart of example steps for replacing image information in . DETAILED DESCRIPTION

[0015] In the following detailed description, numerous specific details are set forth by way of example in order to provide a thorough understanding of the relevant teachings. However, it will be apparent to those skilled in the art that these details are not necessary to practice the present teachings. In other instances, detailed, relatively high-level descriptions of well-known methods, procedures, components, and circuits have been omitted to avoid unnecessarily obscuring aspects of the present teachings.

[0016] As used herein, the term "couple" refers to any logical, optical, physical, electrical connection, link, or the like, whereby signals or light generated or provided by one system element are transferred to another coupled element. Unless otherwise specified, coupled elements or devices need not be in physical contact with each other and may be separated by airspace, intermediate components, elements, or communications media that can modify, manipulate, or carry light or signals.

[0017] The orientations of the eyewear, related components, and any devices shown in any of the figures are exemplary only for purposes of illustration and discussion. In operation, the eyewear may be oriented in other directions, such as upward, downward, sideways, or any other direction, as appropriate for the particular application of the eyewear. Similarly, any directional terms, such as front, back, inward, outward, toward, left, right, lateral, longitudinal, up, down, top, bottom, and side, are exemplary only with respect to directions or orientations and are not intended to be limiting.

[0018] Figure 1A A front perspective view of example eyeglasses 12 for capturing images is depicted. Eyeglasses 12 are shown to include a support structure 13 having temples 14A and 14B extending from a central frame portion 16. Eyeglasses 12 also include hinged joints 18A and 18B, electronic components 20A and 20B, and cores 22A, 22B, and 24. Although shown as eyeglasses, the eyeglasses may take other forms, such as a headset, headgear, helmet, or other device that a user may wear.

[0019] When worn by a user, the support structure 13 supports one or more optical elements within the user's field of view. For example, the central frame portion 16 supports one or more optical elements. As used herein, the term "optical element" refers to a lens, a transparent piece of glass or plastic, a projector, a screen, a display, and other devices for presenting visual images or for a user to perceive visual images. In one example, respective temples 14A and 14B are connected to the central frame portion 16 at respective hinged joints 18A and 18B. The temples 14A and 14B shown are elongated members having cores 22A and 22B extending longitudinally therein.

[0020] exist Figure 1A In FIG. 1 , the temple 14A is shown in a wearable state, and the temple 14B is shown in a folded state. Figure 1A As shown, hinged joint 18A connects temple 14A to right end 26A of central frame portion 16. Similarly, hinged joint 18B connects temple 14B to left end 26B of central frame portion 16. Right end 26A of central frame portion 16 includes a housing for housing electronic assembly 20A therein, and left end 26B includes a housing for housing electronic assembly 20B therein.

[0021] A plastic or other material is embedded in a core wire 22A that extends longitudinally from adjacent hinge joint 18A toward the second longitudinal end of temple 14A. Similarly, plastic or other material is embedded in a core wire 22B that extends longitudinally from adjacent hinge joint 18B toward the second longitudinal end of temple 14B. Plastic or other material is additionally embedded in a core wire 24 that extends from right end 26A (terminating at adjacent electronic assembly 20A) to left end 26B (terminating at adjacent electronic assembly 20B).

[0022] The electronic assemblies 20A and 20B are carried by the support structure 13 (e.g., by one or both of the temples 14A, 14B and / or the central frame portion 16). The electronic assemblies 20A and 20B include a power supply, circuits related to power supply and communication, a communication device, a display device, a computer, a memory, modules and / or the like (not shown). The electronic assemblies 20A and 20B can each include a corresponding imager 10A and 10B for capturing images and / or video. In the example shown, the imager 10A is adjacent to the right temple 14A, and the imager 10B is adjacent to the left temple 14B. The imagers 10A and 10B are spaced apart from each other so as to obtain images of scene objects from two different viewpoints for generating a 3D scene.

[0023] The support structure 13 defines an area (e.g., the area 52 (e.g., the area defined by the frame 12 and the temples 14A and 14B)) for accommodating a portion 52 (e.g., a major portion) of the user / wearer's head. Figure 1B )). The defined area is one or more areas encompassing at least a portion of the user's head that are enclosed by, surround, adjacent to, and / or proximate to the support structure when the user wears the glasses 12. In the example shown, imagers 14A and 14B are positioned on the glasses so that when the glasses 12 are worn, imagers 14A and 14B are adjacent to the user's respective eyes, which helps achieve a separation of viewpoints suitable for creating a three-dimensional scene.

[0024] Figure 21 is a block diagram of example electronic components coupled to a display system 135 (e.g., a display of a processing device or other technology for presenting information). The electronic components shown include a controller 100 (e.g., a hardware processor) for controlling various devices in the glasses 12; a wireless module 102 (e.g., Bluetooth™) for facilitating communication between the glasses 12 and a client device (e.g., a personal computing device 50 such as a smartphone); a power supply circuit 104 (e.g., a battery, a filter, etc.) for powering the glasses 12; a memory 106 such as flash memory for storing data (e.g., images, videos, image processing software, etc.); a selector 32; and one or more imagers 10 (two in the example) for capturing one or more images (e.g., pictures or videos). Although the glasses 12 and the personal computing device are shown as separate components, the functionality of the personal computing device can be incorporated into the glasses, thereby enabling the personal computing device and / or the glasses 12 to perform the functions described herein.

[0025] The selector 32 can trigger (e.g., by momentarily pressing a button) the controller 100 of the glasses 12 to capture an image / video. In an example using a single selector 32, the selector can be used in both a setup mode (e.g., entered by holding the selector 32 for a period of time, such as 3 seconds) and an image capture mode (e.g., entered after a period of no contact, such as 5 seconds) to capture an image.

[0026] In one example, the selector 32 can be a physical button on the glasses 12 that, when pressed, sends a user input signal to the controller 100. The controller 100 can interpret pressing a button within a predetermined time period (e.g., three seconds) as a request to switch to a different operating mode (e.g., enter / exit an operating settings mode). In other examples, the selector 32 can be a virtual button on the glasses or another device. In yet another example, the selector can be a voice module that interprets voice commands, or an eye detection module that detects where the eye is focused. The controller 100 can interpret the signal from the selector 32 as a trigger to cycle through the desired recipient of the image by illuminating the LED 35.

[0027] The wireless module 102 can be coupled to a client / personal computing device 50, such as a smartphone, tablet, phablet, laptop, desktop computer, networking device, access point device, or any other such device that can connect to the wireless module 102. Bluetooth, Bluetooth LE, Wi-Fi, Wi-Fi Direct, cellular modem, and near-field communication systems, as well as multiple instances of any of these systems, can be used to implement these connections to enable communication between them. For example, communication between devices can facilitate the transmission of software updates, images, videos, lighting schemes, and / or sounds between the glasses 12 and the client device.

[0028] In addition, the personal computing device 50 can communicate with one or more recipients (e.g., recipient personal computing devices 51) via a network 53. The network 53 can be a cellular network, Wi-Fi, the Internet, or a similar network that allows the personal computing devices to transmit and receive images, such as via text, email, instant messaging, etc. The computing devices 50 / 51 can each include a processor and a display. Suitable processors and displays, which can be configured to perform the other functions described herein, can be found in current generation personal computing devices and smartphones, such as the iPhone 8 available from Apple Inc. of Cupertino, California. TM and the Samsung Galaxy Note 9, available from Samsung Group in Seoul, South Korea. TM .

[0029] The imager 10 for capturing images / video may include digital camera elements such as a charge coupled device, a lens, or any other light capturing element for capturing image data for conversion into electrical signals.

[0030] The controller 100 controls the electronics. For example, the controller 100 includes circuitry that receives signals from the imager 10 and processes these signals into a format suitable for storage in a memory 106 (e.g., flash memory). The controller 100 powers up and boots into operation in a normal operating mode or into a sleep mode. In one example, the controller 100 includes a microprocessor integrated circuit (IC) customized to process sensor data from the imager 10, and volatile memory that the microprocessor uses for operation. The memory can store software code executed by the controller 100.

[0031] Each electronic component requires a power source to operate. The power supply circuit 104 may include a battery, a power converter, and a power distribution circuit (not shown). The battery may be a rechargeable battery, such as a lithium-ion battery or the like. The power converter and the power distribution circuit may include electronic components for filtering and / or converting voltage for supplying power to various electronic devices.

[0032] Figure 3A A flowchart 300 is depicted, illustrating example operations for rendering a 3D scene on a pair of glasses (e.g., glasses 12 of FIG. 1 ) by a processing system (e.g., a processor of the glasses 12 and / or a processor of a computing device remotely connected to the glasses). For ease of explanation, the steps of flowchart 300 are described with reference to the glasses 12 described herein, but those skilled in the art will recognize that other imager configurations, not limited to glasses, may be used to render a 3D scene. Furthermore, it is understood that one or more steps may be omitted, performed by another component, or performed in a different order.

[0033] In step 310, a first two-dimensional image of a scene object is obtained from a first viewpoint, and a second two-dimensional image of the scene object is obtained from a second viewpoint. In one example, first imager 10A of glasses 12 captures the first two-dimensional image of the scene object from the first viewpoint, and second imager 10B of glasses 12 captures the second two-dimensional image of the scene object from the second viewpoint. The captured images are transmitted from imagers 10A and 10B to a processing system for rendering a three-dimensional scene. In one example, controller 100 of glasses 12 obtains the two-dimensional images and renders a three-dimensional scene from the obtained two-dimensional images. In another example, controller 100 receives the two-dimensional images and transmits them to a processing system of remote computing device 50 / 51 for rendering the three-dimensional scene.

[0034] In step 320, the first and second two-dimensional images are reconstructed into a three-dimensional scene of the scene object. The processing system reconstructs the first and second two-dimensional images into a three-dimensional scene of the scene object. The processing system can use stereo vision processing technology (i.e., creating a depth map) and geometric processing technology to create the three-dimensional scene. In one example, the rendered three-dimensional scene includes geometric features (e.g., vertices with x-axis, y-axis and z-axis coordinates) and image information (e.g., color information, depth information and object information). The rendered three-dimensional scene will also include polygonal faces of connected bodies, such as typical triangular faces or quadrilateral faces, which connect vertices to form a textured surface of the scene object. Based on the description herein, those skilled in the art will understand appropriate stereo vision processing technology and geometric processing technology.

[0035] Figure 3B A flowchart depicting exemplary steps for reconstructing the first and second two-dimensional images into a three-dimensional scene of scene objects (step 320; Figure 3A In step 321, a depth map is created from the first and second two-dimensional images. The processing system may create the depth map by processing the first and second two-dimensional images using stereo vision processing techniques. In step 322, a three-dimensional scene is created from the first and second two-dimensional images. The processing system may create the three-dimensional scene by geometrically processing the first and second two-dimensional images and the depth map created in step 321.

[0036] At step 323, regions of the three-dimensional scene with incomplete image information are detected. The processing system may detect regions of the three-dimensional scene with incomplete image information (e.g., lacking color, depth, and / or object information). For example, the processing system may determine incomplete information by examining the shapes of faces comprising the three-dimensional scene and / or information such as confidence values associated with vertices comprising the faces.

[0037] In one example, the processing system processes the faces of the reconstructed three-dimensional scene (step 323a; Figure 3C) to identify groups of consecutive faces that exhibit incomplete information, such as very narrow faces and / or faces with one or more vertices and low confidence values.

[0038] The processing system may identify adjacent regions with degenerate surfaces (e.g., relatively narrow surfaces) as regions with incomplete information (step 323b; Figure 3C Relatively narrow facets can be determined by comparing the angles between adjacent lines of such facets to a threshold value. For example, a facet can be classified as narrow if it includes at least one angle that is less than a threshold value, such as 5 degrees, if the length of one side of the facet is less than 5% of the length of another side of the facet, and / or if one side of the facet has a dimension that is less than a threshold value, such as 1 mm.

[0039] The processing system may also or alternatively identify faces with low confidence values (eg, faces having at least one vertex with a confidence value below a threshold) as regions with incomplete information (step 323c; Figure 3C ). The confidence value of a vertex may be the confidence value of the corresponding pixel determined during stereo processing of the first and second two-dimensional images to create the depth map (as described above).

[0040] The confidence value of the vertex corresponding to the pixel depends on the match / correlation between the first and second two-dimensional images when the pixel was created. If there is a high correlation (e.g., 75% or higher), then there is a relatively high probability that the vertex contains accurate information useful for reconstructing a three-dimensional scene. On the other hand, if the correlation is low (e.g., below 75%), then there is a relatively high probability that the vertex does not contain information accurate enough to be used to reconstruct the three-dimensional scene.

[0041] In step 324, the detected area is reconstructed. In one example, the processing system reconstructs the detected area of the three-dimensional scene using, for example, geometric processing such as that described in step 322 above. When reconstructing the detected area, the processing system can ignore vertices with low confidence values and / or degenerate faces associated therewith. This results in fewer degenerate faces (e.g., relatively narrow faces), if any. Therefore, after this reconstruction step, the faces in the detected area will have different shapes. The faces in the detected area can be removed before reconstruction. In one example, a data structure such as an Indexed-Face-Set for a three-dimensional mesh can be used, which includes an ordered list of all vertices (and their attributes, such as color, texture, etc.) and a face list, where each face points to a vertex index in the vertex list. In this example, a face can be removed by removing it from the face list.

[0042] In step 325, replacement image information is determined and the reconstructed detected regions are corrected. The processing system may determine replacement image information for the reconstructed detected regions and adjust the 3D scene to include the replacement image information in the detected regions when the replacement image information is reconstructed. For example, the processing system may determine the replacement image information for each detected region by blending boundary information from the respective boundaries of each detected region.

[0043] In one example, to determine the replacement image information, the processing system identifies a boundary around each detected region (step 325a; Figure 3D Then, the processing system identifies background information in the detected area (step 325b; Figure 3D ); for example, based on depth information associated with vertices along the boundary. The processing system also identifies foreground information in the detected region (step 325c; Figure 3D ); and for example, also based on depth information. To identify background / foreground information, the processing system can compare the depth information of each vertex with a threshold (e.g., an average of the depth information from all vertices), identify information associated with vertices with a depth greater than the threshold as background information, and identify information associated with vertices with a depth less than the threshold as foreground information. The processing system then fuses information from the boundaries into the respective regions (step 325d; Figure 3D ), thus giving higher weight to background information than foreground information, which causes information to mainly diffuse from background to foreground.

[0044] Return Reference Figure 3A , in step 330 , the 3D scene is rendered from multiple viewpoints. The processing system can render the 3D scene from multiple viewpoints, for example by applying image synthesis techniques to the 3D scene to create a 2D image from each viewpoint. Based on the description herein, one skilled in the art will understand suitable image synthesis techniques.

[0045] At step 340, the rendered three-dimensional scene is optimized. The processing system can refine the three-dimensional scene from each of the plurality of viewpoints. In one example, the processing system identifies areas of the rendered two-dimensional image of the three-dimensional scene where there are gaps in the image information (i.e., "holes"), and then fills the holes with replacement image information surrounding the holes. The processing system can fill the holes, prioritizing background information surrounding the holes.

[0046] The rendered 3D scene is presented at step 350. The processing system may present the rendered 3D scene on a display of the eyewear or remote computing device by selectively presenting a 2D image within the rendered 3D scene associated with a selected viewpoint (e.g., based on user input to the eyewear device or remote computing device).

[0047] By performing the above process described with reference to flowchart 300, a more aesthetically pleasing 3D scene can be obtained from only two 2D images, viewable from more viewpoints (e.g., simplified or color-stretched), without resorting to panoramic views. Thus, superior results can be achieved without resorting to computationally intensive techniques.

[0048] It should be understood that the steps of the processes described herein can be performed by a hardware processor after loading and executing software code or instructions tangibly stored on a physical computer-readable medium (e.g., magnetic media, a computer hard drive, an optical disk, a solid-state memory, a flash memory, or other storage medium known in the art). Thus, any function performed by a processor described herein can be implemented as software code or instructions tangibly stored on a physical computer-readable medium. After the processor loads and executes such software code or instructions, the processor can perform any function described herein, including any step of the method described herein.

[0049] As used herein, the term "software code" or "code" refers to any instructions or sets of instructions that affect the operation of a computer or controller. They may exist in a computer-executable form (e.g., machine code), which is a set of instructions and data that is directly executed by the computer's central processing unit or controller, in a human-understandable form (e.g., source code) that is executed by the computer's central processing unit or controller, or in an intermediate form (e.g., object code) generated by a compiler. As used herein, the term "software code" or "code" also includes any human-understandable computer instructions or sets of instructions, such as scripts, that can be executed on the fly with the help of an interpreter by a computer's central processing unit or processor.

[0050] Although an overview of the subject matter of the present invention has been described with reference to specific embodiments, various modifications and changes may be made to these examples without departing from the broader scope of the examples of the present disclosure. For example, although the description focuses on eyewear devices, other electronic devices (such as headphones) are also considered to be within the scope of the present subject matter. For convenience only, the term "invention" may be used herein to refer to such examples of the inventive subject matter, individually or collectively, and if there are more than one invention, it is not intended that the scope of this application be automatically limited to any single disclosure or inventive concept, in fact disclosed.

[0051] The examples shown herein are described in sufficient detail to enable those skilled in the art to practice the disclosed teachings. Other examples may be used and derived therefrom, and structural and logical substitutions and changes may be made without departing from the scope of the present disclosure. Therefore, the detailed description should not be read in a limiting sense, and the scope of the various examples is limited only by the appended claims and the full scope of equivalents to which such claims are entitled.

Claims

1. A system for creating a three-dimensional (3D) scene, the system comprising: Eyewear comprising a first imager and a second imager spaced apart from the first imager, the first imager being configured to obtain a first two-dimensional (2D) image of a scene object from a first viewpoint, and the second imager being configured to obtain a second 2D image of the scene object from a second viewpoint different from the first viewpoint; a processing system coupled to the glasses, the processing system being configured to; obtaining the first two-dimensional image and the second two-dimensional image; creating a three-dimensional scene based on the first and second two-dimensional images and connecting vertices within the three-dimensional scene to form a surface; detecting regions of the initial three-dimensional scene having incomplete image information by identifying continuous faces having at least one edge that is smaller than another edge by a predetermined amount or below a threshold size; reconstructing the detected region in the three-dimensional scene; Determine replacement information and correct reconstruction areas; and Rendering a 3D scene with a modified reconstruction region from multiple viewpoints.

2. The system according to claim 1, wherein: The processing system is further configured to: identifying holes in a rendered three-dimensional scene from one or more viewpoints; and The rendered three-dimensional scene is optimized to fill the hole.

3. The system according to claim 1, wherein: The eyeglasses include a first temple and a second temple, and wherein the first imager is adjacent to the first temple and the second imager is adjacent to the second temple.

4. The system according to claim 1, wherein: In order to determine replacement information for the detected region having incomplete image information, the processing system is configured to: identifying a boundary around each detected region; identifying contextual information within the boundaries; identifying foreground information within the boundary; and The background and foreground boundary information of each region is fused so that the weight of the background boundary information is higher than the weight of the foreground boundary information.

5. The system according to claim 4, wherein: To blend the background and foreground boundary information, the processing system is configured to: The missing information of the background boundary information is diffused into the foreground boundary information through each region.

6. The system according to claim 1, wherein: The depth map includes the vertices and corresponding image information from the first and second two-dimensional images, and to reconstruct the detected region, the processing system is configured to connect the vertices of the boundary region to form a second surface, the second surface being different from the first surface.

7. The system according to claim 6, wherein: The first and second faces include at least one of a triangular face or a quadrilateral face.

8. The system according to claim 1, wherein: The incomplete image information also includes degenerate surfaces or low-confidence surfaces.

9. The system according to claim 8, wherein: Each low-confidence surface includes at least one vertex generated with incompatible values between the first two-dimensional image and the second two-dimensional image.

10. A method for creating a three-dimensional (3D) scene, the method comprising: obtaining a first two-dimensional (2D) image of a scene object from a first viewpoint; obtaining a second two-dimensional image of the scene object from a second viewpoint different from the first viewpoint; creating a three-dimensional scene based on the first and second two-dimensional images and connecting vertices within the three-dimensional scene to form a surface; detecting regions of the initial three-dimensional scene having incomplete image information by identifying continuous faces having at least one edge that is smaller than another edge by a predetermined amount or below a threshold size; reconstructing the detected region in the three-dimensional scene; Determine replacement information and correct reconstruction areas; and Rendering a 3D scene with a modified reconstruction region from multiple viewpoints.

11. The method according to claim 10, further comprising: identifying holes in a rendered three-dimensional scene from one or more viewpoints; and The rendered three-dimensional scene is optimized to fill the hole.

12. The method according to claim 10, wherein: The first two-dimensional image is obtained from a first imager adjacent to a first temple of the eyeglasses, and the second two-dimensional image is obtained from a second imager adjacent to a second temple of the eyeglasses.

13. The method according to claim 10, wherein: The determining step comprises: identifying a boundary around each detected region; identifying contextual information within the boundaries; identifying foreground information within the boundary; and The background and foreground boundary information of each region is fused so that the weight of the background boundary information is higher than the weight of the foreground boundary information.

14. The method according to claim 13, wherein The mixing includes: The missing information of the background boundary information is diffused into the foreground boundary information through each region.

15. The method according to claim 10, wherein The step of creating a three-dimensional scene comprises creating a depth map comprising the vertices and corresponding image information from the first and second two-dimensional images, and wherein the reconstructing step further comprises: The vertices of the boundary region are connected to form a second face, wherein the second face is different from the first face.

16. The method according to claim 15, wherein The first and second faces include at least one of a triangular face or a quadrilateral face.

17. The method according to claim 10, wherein The incomplete image information also includes degenerate surfaces or low-confidence surfaces.

18. The method according to claim 17, wherein Each low-confidence surface includes at least one vertex generated by an incompatibility value exceeding a threshold between the first two-dimensional image and the second two-dimensional image.