Method and electronic device for augmented reality
By using depth change reprojection technology in the video perspective XR system, the foreground objects are accurately projected, and the background objects are constant depth projected, which solves the problems of high computing resources and delays and improves system performance and user experience.
Patent Information
- Application Number
- CN202380078524.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2023-12-01
- Filing Date
- 2023-12-27
- Publication Date
- 2025-07-08
AI Technical Summary
Existing video perspective (VST) XR systems have high computing resources demands and prominent latency problems during depth-based reprojection, resulting in poor user experience.
Depth change reprojection technology is used to accurately project the depth reprojection by identifying the foreground objects in the scene, while background objects adopt constant depth projection, reducing calculation load and improving efficiency.
It significantly reduces the computing resources and time requirements during the deep reconstruction process, reduces latency, and improves the performance and user experience of the VST XR system.
Smart Images

Figure CN120283267A_ABST
Abstract
Description
Technical Field
[0001] The present disclosure generally relates to extended reality (XR) systems and processes. More specifically, the present disclosure relates to depth change reprojection penetration in video see-through (VST) XR. Background Art
[0002] Over time, extended reality (XR) systems have become increasingly popular, and many applications have been or are being developed for XR systems. Some XR systems, such as augmented reality (AR) systems and mixed reality (MR) systems, can enhance a user's view of their current environment by overlaying digital content, such as information or virtual objects, on the user's view of the current environment. For example, some XR systems can generally seamlessly blend computer-graphics-generated virtual objects with real-world scenes. Summary of the Invention
[0003] Technical Solution
[0004] The present disclosure relates to depth change reprojection penetration in video see-through (VST) extended reality (XR).
[0005] In an embodiment, a method includes: obtaining (i) an image of a captured scene using a stereoscopic imaging sensor of an electronic device, and (ii) depth data associated with the image, wherein the scene includes a plurality of objects. The method further includes: obtaining a volume-based three-dimensional (3D) model of the plurality of objects included in the scene. The method further includes: for one or more first objects of the plurality of objects, performing depth-based reprojection of one or more 3D models of the one or more first objects to a left virtual view and a right virtual view based on one or more depths of the one or more first objects. The method further includes: for one or more second objects of the plurality of objects, performing constant-depth reprojection of one or more 3D models of the one or more second objects to the left virtual view and the right virtual view based on a specified depth. Additionally, the method further includes rendering the left virtual view and the right virtual view for presentation by the electronic device.
[0006] In an embodiment, an electronic device includes at least one display. The electronic device also includes an imaging sensor configured to capture an image of a scene, where the scene includes a plurality of objects. The electronic device further includes: at least one processor configured to: obtain (i) an image of the scene and (ii) depth data associated with the image, and obtain a volume-based 3D model of the plurality of objects included in the scene. The at least one processor is further configured to, for one or more first objects among the plurality of objects, perform depth-based reprojection of one or more 3D models of the one or more first objects to a left virtual view and a right virtual view based on one or more depths of the one or more first objects. The at least one processor is further configured to, for one or more second objects among the plurality of objects, perform constant-depth reprojection of one or more 3D models of the one or more second objects to a left virtual view and a right virtual view based on a specified depth. Additionally, the at least one processor is configured to render the left virtual view and the right virtual view for presentation by the at least one display.
[0007] In an embodiment, a machine-readable medium includes instructions that, when executed, cause at least one processor to: obtain (i) an image of a scene captured using a stereoscopic imaging sensor of an electronic device; and (ii) depth data associated with the image, where the scene includes a plurality of objects. The machine-readable medium further includes instructions that, when executed, cause at least one processor to: obtain a volume-based 3D model of the plurality of objects included in the scene. The machine-readable medium further includes instructions that, when executed, cause at least one processor to: for one or more first objects among the plurality of objects, perform depth-based reprojection of one or more 3D models of the one or more first objects to a left virtual view and a right virtual view based on one or more depths of the one or more first objects. The machine-readable medium further includes instructions that, when executed, cause at least one processor to: for one or more second objects among the plurality of objects, perform constant-depth reprojection of one or more 3D models of the one or more second objects to a left virtual view and a right virtual view based on a specified depth. Additionally, the machine-readable medium includes instructions that, when executed, cause at least one processor to render the left virtual view and the right virtual view for presentation by the electronic device.
[0008] For those skilled in the art, other technical features can be readily apparent from the following drawings, description, and claims. BRIEF DESCRIPTION OF THE DRAWINGS
[0009] To more fully understand the present disclosure and its advantages, reference is now made to the following description taken in conjunction with the accompanying drawings, in which:
[0010] Figure 1 An example network configuration including an electronic device in accordance with the present disclosure is shown;
[0011] Figure 2 Shows an example process for reprojecting penetration of depth changes in video see-through (VST) extended reality (XR) according to the present disclosure;
[0012] Figure 3 Shows an example pipeline for reprojecting penetration of depth changes in VST XR according to the present disclosure;
[0013] Figure 4 Shows a specific example implementation of a pipeline for reprojecting penetration of depth changes in VST XR according to the present disclosure;
[0014] Figure 5 Shows an example process for three-dimensional (3D) model reconstruction using image-guided depth fusion according to the present disclosure;
[0015] Figure 6 Shows an example consistency between the left and right images of a stereoscopic image pair according to the present disclosure;
[0016] Figure 7 Shows an example process for separating a scene into foreground and background objects based on user focus according to the present disclosure;
[0017] Figure 8 Shows an example process for separating a scene into foreground and background objects using machine learning according to the present disclosure;
[0018] Figure 9 Shows an example of depth-based reprojection for generating left and right virtual views according to the present disclosure;
[0019] Figure 10 Shows an example of depth change reprojection for selected objects at different depth levels according to the present disclosure;
[0020] Figure 11 Shows an example process for building a library of 3D reconstructed scenes and objects according to the present disclosure;
[0021] Figure 12 Shows an example of generating one virtual image at one viewpoint using another virtual image at another viewpoint according to the present disclosure; and
[0022] Figure 13 Shows an example method for reprojecting penetration of depth changes in VST XR according to the present disclosure. Detailed Description
[0023] It may be beneficial to clarify the definitions of certain words and phrases used throughout this patent document. The terms "send", "receive", and "communicate" and their derivatives cover both direct and indirect communication. The terms "comprise" and "include" and their derivatives mean including but not limiting. The term "or" is inclusive and means and / or. The phrase "associated with" and its derivatives mean including, being included within, interconnected with, containing, being contained within, connected to or connected with, coupled to or coupled with, communicable with, cooperating with, interlaced, juxtaposed, proximate, bound or bound with, having, having the attribute of, having a relationship with, and so on.
[0024] In addition, the various functions described below can be implemented or supported by one or more computer programs, each of which consists of computer-readable program code and is embedded in a computer-readable medium. The terms "application" and "program" refer to one or more computer programs, software components, instruction sets, processes, functions, objects, classes, instances, related data, or a part thereof, which are suitable for implementation with appropriate computer-readable program code. The phrase "computer-readable program code" includes any type of computer code, including source code, object code, and executable code. The phrase "computer-readable medium" includes any type of medium that a computer can access, such as read-only memory (ROM), random access memory (RAM), hard disk drive, compact disc (CD), digital video disc (DVD), or any other type of memory. The computer-readable medium includes a medium that can permanently store data and a medium that can store data and then overwrite it, such as a rewritable optical disc or an erasable storage device.
[0025] Terms and phrases used herein, such as "has", "may have", "include", or "may include" a feature (such as a number, function, operation, or a component such as a part), indicate the presence of the feature and do not exclude the presence of other features. In addition, the phrases "A or B", "at least one of A and / or B", or "one or more of A and / or B" used herein may include all possible combinations of A and B. For example, "A or B", "at least one of A and B", and "at least one of A or B" may indicate all of the following cases: (1) including at least one A; (2) including at least one B; or (3) including at least one A and at least one B. In addition, the terms "first" and "second" used herein may modify various components regardless of their importance and do not limit these components. These terms are only used to distinguish one component from another. For example, a first user device and a second user device may indicate different user devices from each other, regardless of the order or importance of the devices. Without departing from the scope of the present disclosure, the first component may be represented as the second component, and vice versa.
[0026] It should be understood that when an element (e.g., a first element) is referred to as being (operatively or communicatively) "coupled to" or "connected to" another element (e.g., a second element) / "coupling" or "connecting" with another element (e.g., a second element), the element can be coupled or connected to the other element directly or via a third element. Conversely, it should be understood that when an element (e.g., a first element) is referred to as "directly coupled to" or "directly connected to" another element (e.g., a second element) / "directly coupling" or "directly connecting" with another element (e.g., a second element), there is no other element (e.g., a third element) between the element and the other element.
[0027] The phrase "configured (or set) to" used herein can be used interchangeably with the phrases "suitable for", "capable of", "designed to", "adapted to", "manufactured to", or "able to" depending on the context. The phrase "configured (or set) to" does not necessarily mean "specially designed in hardware for...". Instead, the phrase "configured to" can mean that a device can perform operations together with other devices or components. For example, the phrase "a processor configured (or set) to execute A, B, and C" can refer to a general-purpose processor (e.g., a CPU or an application processor) that can perform operations by executing one or more software programs stored in a storage device, or can refer to a dedicated processor (e.g., an embedded processor) for performing operations.
[0028] The terms and phrases used herein are only for describing some embodiments of the present disclosure and are not intended to limit the scope of other embodiments of the present disclosure. It should be understood that unless the context clearly dictates otherwise, the singular forms "a", "an", and "the" also include the plural meanings. All terms and phrases, including technical and scientific terms and phrases, used herein have the same meanings as those commonly understood by those of ordinary skill in the art to which the embodiments of the present disclosure belong. It should also be understood that terms and phrases (e.g., terms and phrases defined in a common dictionary) should be interpreted as having meanings consistent with their meanings in the context of the relevant art, and should not be interpreted in an idealized or overly formal manner unless clearly defined herein otherwise. In some cases, the terms and phrases defined herein may be interpreted as excluding embodiments of the present disclosure.
[0029] Examples of an "electronic device" according to an embodiment of the present disclosure may include at least one of the following: a smart phone, a tablet personal computer (PC), a mobile phone, a video phone, an e-book reader, a desktop PC, a laptop computer, a netbook computer, a workstation, a personal digital assistant (PDA), a portable multimedia player (PMP), an MP3 player, a mobile medical device, a camera, or a wearable device (e.g., smart glasses, a head-mounted device (HMD), electronic clothing, an electronic bracelet, an electronic necklace, electronic accessories, an electronic tattoo, a smart mirror, or a smart watch). Other examples of the electronic device include smart home appliances. Examples of the smart home appliances may include at least one of the following: a television, a digital video disc (DVD) player, an audio player, a refrigerator, an air conditioner, a vacuum cleaner, an oven, a microwave oven, a washing machine, a dryer, an air purifier, a set-top box, a home automation control panel, a security control panel, a TV box (e.g., SAMSUNG HOMESYNC, APPLE TV, or GOOGLE TV), a smart speaker, or a speaker integrated with a digital assistant (e.g., SAMSUNG GALAXY HOME, APPLE HOMEPOD, or AMAZON ECHO), a game console (e.g., XBOX, PLAYSTATION, or NINTENDO), an electronic dictionary, an electronic key, a camera, or an electronic photo frame. Other examples of the electronic device include at least one of various medical devices (e.g., various portable medical measurement devices such as a blood glucose measurement device, a heartbeat measurement device, or a body temperature measurement device, a magnetic resource angiography (MRA) device, a magnetic resource imaging (MRI) device, a computed tomography (CT) device, an imaging device, or an ultrasonic device), a navigation device, a global positioning system (GPS) receiver, an event data recorder (EDR), a flight data recorder (FDR), an in-vehicle infotainment device, marine electronics (e.g., a marine navigation device or a gyrocompass), avionics, a security device, an in-vehicle head unit, an industrial or domestic robot, an automated teller machine (ATM), a point of sale (POS) device, or an Internet of Things (IoT) device (e.g., a light bulb, various sensors, a water meter, a gas meter, a sprinkler, a fire alarm, a thermostat, a street lamp, a toaster, a fitness device, a hot water tank, a heater, or a boiler). Other examples of the electronic device include at least a part of furniture or a building / structure, an electronic whiteboard, an electronic signature receiving device, a projector, or various measurement devices (e.g., devices for measuring water, electricity, gas, or electromagnetic waves). It should be noted that according to various embodiments of the present disclosure, the electronic device may be one or a combination of the above devices. According to some embodiments of the present disclosure, the electronic device may be a flexible electronic device. The electronic device disclosed herein is not limited to the above devices and may further include any other currently known or subsequently developed electronic device.
[0030] In the following description, an electronic device will be described with reference to the accompanying drawings in accordance with various embodiments of the present disclosure. The term "user" as used herein may refer to a person or another device (e.g., an artificial intelligence electronic device) that uses the electronic device.
[0031] Throughout this patent document, definitions of other specific words and phrases may be provided. Those skilled in the art should understand that in many, if not most, cases, these definitions apply to both prior and future uses of the words and phrases so defined.
[0032] A description will be given below with reference to the accompanying drawings Figures 1 to 13 and various embodiments of the present disclosure. However, it should be understood that the present disclosure is not limited to these embodiments, and all changes and / or equivalents or alternatives thereof also fall within the scope of the present disclosure. Throughout the specification and the drawings, the same or similar reference numerals may be used to refer to the same or similar elements.
[0033] As described above, extended reality (XR) systems are becoming increasingly popular, and numerous applications have been and are being developed for XR systems. Some XR systems (e.g., augmented reality (AR) systems and mixed reality (MR) systems) can enhance a user's view of their current environment by overlaying digital content (e.g., information or virtual objects) on the user's view of the current environment. For example, some XR systems can typically seamlessly blend computer-graphics-generated virtual objects with real-world scenes.
[0034] An optical see-through (OST) XR system is an XR system in which a user directly views a real-world scene through a head-mounted device (HMD). Unfortunately, OST XR systems face numerous challenges that may limit their applications. These challenges include: limited field of view, limited usage space (e.g., only indoor use), inability to display completely opaque black objects, and the use of a complex optical pipeline that may require projectors, waveguides, and other optical elements. In contrast to OST XR systems, a video see-through (VST) XR system (also referred to as a "passthrough" XR system) presents a generated video sequence of a real-world scene to the user. A VST XR system can be built using virtual reality (VR) technology and has various advantages over OST XR systems. For example, a VST XR system can provide a wider field of view and can provide improved context-aware augmented reality (AR).
[0035] Viewpoint matching is generally a useful or important operation in a VST XR pipeline. Viewpoint matching generally refers to the process of creating a video frame that is presented at a user's eye viewpoint position using a video frame captured at a perspective camera viewpoint position, which makes the user feel as if the perspective camera is located at the user's eye viewpoint position. Additionally, viewpoint matching can also involve depth-based reprojection, in which objects in a scene are reprojected into a virtual view based on their depth in the scene. However, depth-based reprojection may require a large amount of computational resources (such as processing and memory resources) to reconstruct depth and perform depth reprojection, which is particularly problematic at higher video resolutions (such as 4K and above). Additionally, depth-based reprojection may introduce latency in the VST XR pipeline, resulting in noticeable delays or other issues for the user.
[0036] The present disclosure provides a depth-varying reprojection penetration technique in VST XR. As described in more detail below, an image of a scene captured using a stereoscopic imaging sensor pair of an XR device and depth data associated with these images are obtained, where the scene includes a plurality of objects. Additionally, a volumetric three-dimensional (3D) model of the plurality of objects included in the scene is obtained. For example, the XR device can determine whether there is a volumetric 3D model corresponding to an object in a previously generated library of volumetric 3D models for each object in the scene. For example, a volumetric 3D model can refer to a three-dimensional representation of an object or a scene defined by its volume or physical boundaries. If there is, the corresponding volumetric 3D model in the library can be used for the object. If not, a new volumetric 3D model can be generated for the object and stored in the library for future use.
[0037] For one or more first objects among a plurality of objects, depth-based reprojection of one or more 3D models of the one or more first objects to a left virtual view and a right virtual view can be performed based on one or more depths of the one or more first objects. For one or more second objects among the plurality of objects, constant-depth reprojection of one or more 3D models of the one or more second objects to the left virtual view and the right virtual view can be performed based on a specified depth. In some cases, constant-depth reprojection can be performed using a shader of a graphics processing unit (GPU), in which case the computational load of the constant-depth reprojection can be very small. The left virtual view and the right virtual view can be rendered and presented by an XR device. The first objects and the second objects can be identified in various ways (such as by using a trained machine learning model that segments the scene or by identifying a focus region at which the user appears to be gazing). In some cases, one or more first objects can include one or more foreground objects in the scene, and one or more second objects can include one or more background objects in the scene.
[0038] In this way, these techniques provide an efficient mechanism for depth-varying reprojection, which can be used to support functions such as view-point matching and head-pose change compensation. The techniques can identify the exact depth of one or more first objects and can use a constant depth for one or more second objects. This enables geometric reconstruction of the one or more first objects at an exact depth, while the one or more second objects can be projected onto a plane with a constant depth. Thus, a final virtual view can be generated by performing exact depth-based reprojection for some objects and constant-depth reprojection for other objects. These techniques help significantly reduce computational resources and computational time in the depth reconstruction process because only some objects may require exact depth and the depth of other objects can be set to a constant value. Additionally, since only one or more first objects can be depth-based reprojected and since 3D models are used, all 3D information of the one or more first objects is available, and thus almost no hole artifacts are generated by the depth reprojection performed herein. This can further simplify the depth reprojection process because hole artifacts may not need to be removed or the degree of removal required is much smaller. Overall, these techniques can help significantly improve the performance of the VST XR pipeline.
[0039] Figure 1 An example network configuration 100 including an electronic device is shown in accordance with the present disclosure. Figure 1 The embodiments of the network configuration 100 shown are for reference only. Other embodiments of the network configuration 100 can be used without departing from the scope of the present disclosure.
[0040] According to an embodiment of the present disclosure, an electronic device 101 is included in the network configuration 100. The electronic device 101 may include at least one of a bus 110, a processor 120, a memory 130, an input / output (I / O) interface 150, a display 160, a communication interface 170, and a sensor 180. In some embodiments, the electronic device 101 may exclude at least one of these components, or may add at least one other component. The bus 110 includes circuitry for connecting the components 120-180 to each other and transmitting communications (e.g., control messages and / or data) between the components.
[0041] The processor 120 includes one or more processing devices, such as one or more microprocessors, microcontrollers, digital signal processors (DSPs), application specific integrated circuits (ASICs), or field programmable gate arrays (FPGAs). In some embodiments, the processor 120 includes one or more of a central processing unit (CPU), an application processor (AP), a communication processor (CP), a graphics processing unit (GPU), or a neural processing unit (NPU). The processor 120 is capable of controlling at least one other component of the electronic device 101, and / or performing operations or data processing related to communications or other functions. As described below, the processor 120 may perform one or more functions related to depth change reprojection penetration in VST XR.
[0042] The memory 130 may include volatile memory and / or non-volatile memory. For example, the memory 130 may store commands or data related to at least one other component of the electronic device 101. According to an embodiment of the present disclosure, the memory 130 may store software and / or a program 140. The program 140 includes, for example, a kernel 141, middleware 143, an application programming interface (API) 145, and / or an application program (or “app”) 147. At least a portion of the kernel 141, middleware 143, or API 145 may be represented as an operating system (OS).
[0043] The kernel 141 can control or manage system resources (such as the bus 110, the processor 120, or the memory 130) for performing operations or functions implemented in other programs (such as middleware 143, API 145, or application 147). The kernel 141 provides an interface that allows the middleware 143, API 145, or application 147 to access the various components of the electronic device 101 to control or manage system resources. The application 147 may include one or more applications that, among other functions, perform depth-varying reprojection penetration in VST XR. These functions may be performed by a single application or may be performed by multiple applications, each performing one or more of these functions. For example, the middleware 143 may act as a repeater, allowing the API 145 or application 147 to communicate data with the kernel 141. Multiple applications 147 may be provided. The middleware 143 is capable of controlling work requests received from the application 147, for example, by assigning priorities for using the system resources of the electronic device 101 (such as the bus 110, the processor 120, or the memory 130) to at least one of the multiple applications 147. The API 145 is an interface that allows the application 147 to control the functions provided by the kernel 141 or the middleware 143. For example, the API 145 includes at least one interface or function (such as a command) for file archiving control, window control, image processing, or text control.
[0044] The I / O interface 150 serves as an interface that, for example, can send commands or data input from a user or other external device to other components of the electronic device 101. The I / O interface 150 can also output commands or data received from other components of the electronic device 101 to the user or other external device.
[0045] The display 160 includes, for example, a liquid crystal display (LCD), a light emitting diode (LED) display, an organic light emitting diode (OLED) display, a quantum dot light emitting diode (QLED) display, a microelectromechanical systems (MEMS) display, or an electronic paper display. The display 160 may also be a depth perception display, such as a multi-focus display. The display 160 is capable of displaying various contents (such as text, images, videos, icons, or symbols) to the user. The display 160 may include a touch screen and may receive, for example, touch, gesture, proximity, or hover inputs using an electronic pen or a user's body part.
[0046] The communication interface 170 is capable of establishing communication, for example, between the electronic device 101 and an external electronic device (such as the first electronic device 102, the second electronic device 104, or the server 106). For example, the communication interface 170 may be connected to the network 162 or 164 through a wireless or wired communication to communicate with the external electronic device. The communication interface 170 may be a wired or wireless transceiver, or any other component for transmitting and receiving signals.
[0047] Wireless communication can use at least one of, for example, WiFi, Long Term Evolution (LTE), Long Term Evolution-Advanced (LTE-A), Fifth Generation Wireless System (5G), millimeter wave or 60 GHz wireless communication, Wireless USB, Code Division Multiple Access (CDMA), Wideband Code Division Multiple Access (WCDMA), Universal Mobile Telecommunications System (UMTS), Wireless Broadband (WiBro), or Global System for Mobile Communications (GSM) as a communication protocol. Wired connections can include at least one of, for example, Universal Serial Bus (USB), High-Definition Multimedia Interface (HDMI), Recommended Standard 232 (RS-232), or Plain Old Telephone Service (POTS). Network 162 or 164 includes at least one communication network, such as a computer network (e.g., Local Area Network (LAN) or Wide Area Network (WAN)), the Internet, or a telephone network.
[0048] The electronic device 101 also includes: one or more sensors 180 for measuring a physical quantity or detecting an activation state of the electronic device 101 and converting the measured or detected information into an electrical signal. For example, the sensor 180 includes: a camera or other imaging sensors that can be used to capture an image of a scene. The sensor 180 can also include one or more buttons for touch input, one or more microphones, a depth sensor, a gesture sensor, a gyroscope or gyroscopic sensor, a barometric pressure sensor, a magnetic sensor or magnetometer, an acceleration sensor or accelerometer, a grip sensor, a proximity sensor, a color sensor (e.g., Red Green Blue (RGB) sensor), a biophysical sensor, a temperature sensor, a humidity sensor, an illuminance sensor, an ultraviolet (UV) sensor, an electromyogram (EMG) sensor, an electroencephalogram (EEG) sensor, an electrocardiogram (ECG) sensor, an infrared (IR) sensor, an ultrasonic sensor, an iris sensor, or a fingerprint sensor. Additionally, the sensor 180 can include one or more position sensors, such as an inertial measurement unit that can include one or more accelerometers, gyroscopes, and other components. Further, the sensor 180 can include control circuitry for controlling at least one sensor included herein. Any of these sensors 180 can be located within the electronic device 101.
[0049] In some embodiments, the electronic device 101 may be a wearable device or a wearable device that can be mounted on an electronic device (such as an HMD). For example, the electronic device 101 may represent an XR wearable device, such as earphones or smart glasses. In some embodiments, the first external electronic device 102 or the second external electronic device 104 may be a wearable device or a wearable device that can be mounted on an electronic device (such as an HMD). In some embodiments, when the electronic device 101 is mounted in the electronic device 102 (such as an HMD), the electronic device 101 may communicate with the electronic device 102 through the communication interface 170. The electronic device 101 may be directly connected to the electronic device 102 to communicate with the electronic device 102 without involving a separate network.
[0050] The first external electronic device 102, the second external electronic device 104, and the server 106 may each be a device of the same or different type as the electronic device 101. According to certain embodiments of the present disclosure, the server 106 includes a group of one or more servers. In addition, according to certain embodiments of the present disclosure, all or part of the operations performed on the electronic device 101 may be performed on one or more other electronic devices (such as the electronic devices 102 and 104 or the server 106). In addition, according to certain embodiments of the present disclosure, when the electronic device 101 needs to automatically or upon request perform some functions or services, the electronic device 101 may request another device (such as the electronic devices 102 and 104 or the server 106) to perform at least some functions related thereto, rather than performing the function or service itself or additionally. The other electronic devices (such as the electronic devices 102 and 104 or the server 106) are capable of performing the requested function or additional functions and sending the execution results to the electronic device 101. The electronic device 101 may provide the requested function or service by processing the received results as they are or additionally. For this purpose, for example, cloud computing, distributed computing, or client-server computing technologies may be used. Although Figure 1 it is shown that the electronic device 101 includes a communication interface 170 to communicate with the external electronic device 104 or the server 106 through the network 162 or 164, according to some embodiments of the present disclosure, the electronic device 101 may operate independently without a separate communication function.
[0051] The server 106 may include components (or a suitable subset thereof) that are the same as or similar to those of the electronic device 101. The server 106 may support driving the electronic device 101 by executing at least one operation (or function) implemented on the electronic device 101. For example, the server 106 may include a processing module or a processor that may support the processor 120 implemented in the electronic device 101. As described below, the server 106 may perform one or more functions related to depth change reprojection penetration in VST XR.
[0052] Although Figure 1 an example of a network configuration 100 including an electronic device 101 is shown, various changes can be made to Figure 1 it. For example, the network configuration 100 can include any number of various components in any suitable arrangement. Generally, computing and communication systems have a wide variety of configurations, and Figure 1 the scope of the present disclosure is not limited to any particular configuration. Additionally, although Figure 1 an operating environment in which various features disclosed in this patent document can be used is shown, these features can also be used in any other suitable system. Hereinafter, the electronic device 101 can be referred to as, for example, an XR device.
[0053] Figure 2 An example process 200 for implementing depth change reprojection penetration in VST XR according to the present disclosure is shown. For ease of explanation, Figure 2 the process 200 in Figure 1 is described as being executed or implemented by the electronic device 101 in the network configuration 100 of
[0054] As Figure 2 shown, the process 200 includes an image capture operation 202 in which an image of a scene is captured or otherwise obtained. For example, the image of the captured scene can be captured using the stereoscopic imaging sensor of the XR device, such as capturing an image of 180 using the stereoscopic imaging sensor of the electronic device 101. As described below, depth information associated with the scene can also be obtained. The depth information can be obtained in any suitable manner (e.g., by using at least one depth sensor 180 of the electronic device 101). Using the captured image and the depth information, a scene model reconstruction operation 204 can be performed to identify the 3D model of the overall scene captured in the image and the 3D models of the individual objects detected within the scene captured in the image. As described below, one or more 3D models can optionally be retrieved from a 3D model library.
[0055] A background and foreground separation operation 206 can be performed to separate the reconstructed 3D scene into foreground and background. In some embodiments, the depth information associated with the objects detected in the scene can be used to separate the reconstructed 3D scene into foreground and background. For example, objects within a threshold distance from the XR device can be considered foreground objects, while objects outside the threshold distance from the XR device can be considered background objects. In certain embodiments, a trained machine learning model can be used to automatically determine the threshold distance. In some embodiments, separating the reconstructed 3D scene into foreground and background can be accomplished by tracking the region that the user is gazing at within the scene and considering that region as the user's focus region. Objects within the focus region can be considered foreground objects, while objects outside the focus region can be considered background objects. It should be noted that any other suitable techniques can be used herein to separate the reconstructed 3D scene into foreground and background.
[0056] Based on the 3D models of the objects and the scene and the segmentation of the scene into foreground and background, a virtual view generation operation 208 can be performed to generate virtual views of the scene for presentation to the user of the XR device. Herein, the virtual view generation operation 208 can generate entirely new views of the scene based on factors such as the user's head pose. The virtual view generation operation 208 can generate left and right views for presentation by the XR device, for example, via presentation on one or more displays 160 of the electronic device 101. As part of the virtual view generation operation 208, depth-based reprojection (possibly with higher accuracy) can be performed on the identified foreground objects based on the actual depth of those foreground objects in the scene. A simpler constant-depth reprojection (possibly with lower accuracy) can be performed on the identified background objects based on a specified constant depth in the scene.
[0057] Generally, operations 202 - 208 of process 200 can be used together to implement the depth - change reprojection operation 210. The depth - change reprojection operation 210 herein is used to convert video frames captured at multiple perspective - imaging sensor viewpoint positions 212 (each position may have up to six degrees of freedom in possible position and orientation) into video frames that appear to be captured at multiple eye - viewpoint positions 214 of the user (each position may have up to six degrees of freedom in possible position and orientation). Ideally, from the user's perspective, the images presented to the user seem to be captured by the imaging sensor 180 located at the user's eye position, while in fact, these images are captured using imaging sensors 180 located at other positions. The reprojection operation 210 is referred to as "depth - change" herein because the reprojection operation 210 can include (i) depth - based reprojection of objects and (ii) constant - depth reprojection of other objects. "Depth - based" reprojection refers to the reprojection of the 3D model of an object based on the actual depth of the object in the scene. "Constant - based" reprojection refers to projecting the 3D model of an object to a specified constant depth regardless of the actual depth of the object in the scene.
[0058] Although Figure 2 shows an example of the depth - change reprojection penetration process 200 in VST XR, various changes can be made Figure 2 For example, various functions in Figure 2 can be combined, further subdivided, replicated, omitted, or rearranged, and additional functions can be added according to specific needs.
[0059] Figure 3 shows an example pipeline 300 for depth - change reprojection penetration in VST XR according to the present disclosure. For ease of illustration, Figure 3 the pipeline 300 in Figure 1 is described as being implemented using the electronic device 101 in the network configuration 100 of Figure 2 where the pipeline 300 can be used to perform at least a part of the process 200 of
[0060] As Figure 3As shown, pipeline 300 includes various modules that represent components or functions for implementing the various operations of the above-described process 200. In this example, pipeline 300 includes a sensor module 302 and a data capture module 304. Sensor module 302 may include various sensors 180 used by the XR device, and data capture module 304 may be used to capture or otherwise obtain various data from the sensors of sensor module 302. For example, sensor module 302 may include multiple imaging sensors, at least one depth sensor, and at least one position sensor. The imaging sensors may be used to capture scene images. The at least one depth sensor may be used to obtain depth information associated with the scene, such as in the form of a depth map. A depth map may represent a mapping associated with at least one image that estimates the depth within the scene captured in the corresponding image. In some cases, for example, when a signal is sent from and received at the at least one depth sensor to support time-of-flight (ToF) measurements, the at least one depth sensor may use an active technique to identify depth. The at least one position sensor may be used to identify the position and orientation of the XR device in a given environment. In some cases, the at least one position sensor may include a pose tracking camera for capturing images and an IMU for sensing the orientation of the XR device. The data capture module 304 herein may be used to obtain various information, such as captured scene images, depth maps, or other depth information associated with the scene, as well as the pose of the XR device.
[0061] A data fusion module 306 processes the acquired data to fuse the depth data. For example, data fusion module 306 may be used to fuse the depth from the depth sensors, the depth measured using the pose sensors, and other depth data from other sensors. Data fusion module 306 may also be used to clarify and reconstruct the depth map of the captured scene. For example, data fusion module 306 may identify locations where the identified depth may be inaccurate / inconsistent or missing and replace these depths with other depths (such as the average depth of surrounding pixels). Data fusion module 306 may also generate a fused depth map based on the received information, where the fused depth map represents an initial estimate of the depth within the scene based on the information received from various sensors.
[0062] The depth densification module 308 processes the fused depth map to generate a higher resolution depth map. For example, the depth densification module 308 can perform depth densification and super-resolution processing, which increases the number of depth values in the higher resolution depth map compared to the fused depth map. As shown, in some cases, the depth densification module 308 can be guided at least in part by the library 310, which represents a collection of stored 3D models for reconstructing scenes and reconstructed objects within the scene. The stored 3D models can be associated with scenes and objects previously encountered by the XR device, or with scenes and objects that the XR device has been preconfigured to recognize. Since the depth densification module 308 can use the 3D models to help increase the density of depth values, the depth densification module 308 can use the 3D model of the scene and / or the 3D models of one or more objects when performing depth densification.
[0063] The scene model calculation and reconstruction module 312 processes the higher resolution depth map and obtains one or more 3D models from the library 310 (if available) in order to generate or otherwise obtain 3D models for the captured scene and for the objects within the captured scene. For example, the scene model calculation and reconstruction module 312 can generate or otherwise obtain a 3D model of the actual environment surrounding the XR device, as well as a 3D model of each object detected in the surrounding environment of the XR device. In some cases, the 3D model of the environment and / or the 3D models of one or more objects can be retrieved from the library 310 and used by the scene model calculation and reconstruction module 312. If the scene or an object within the scene cannot be recognized (meaning that the 3D model is missing from the library 310), then the scene model calculation and reconstruction module 312 can generate a new 3D model for the scene or object and store the new 3D model in the library 310. In some embodiments, each 3D model in the library 310 can include a 3D mesh representing the geometry of the associated scene or object, a color texture map representing at least one color of the associated scene or object, and one or more parameters for identifying at least one of a pattern, size, pose, and transformation of the associated scene or object. In this document, these 3D models can be referred to as "volume-based" 3D models because they are volume-based 3D object representations.
[0064] The virtual view generation module 314 reprojects the 3D model of the scene and the objects within the scene to generate a virtual view of the scene. As described above and in more detail below, when generating the virtual view, the virtual view generation module 314 can perform depth-based reprojection on certain objects in the scene and perform constant-depth reprojection on other objects in the scene. The information blending module 316 can perform the blending of real-world information (which can be captured using a perspective imaging sensor) and virtual information (which can be generated by a graphics pipeline or other components). For example, this can enable the information blending module 316 to combine the representation of real-world objects in the scene with digital content (such as one or more virtual objects) generated by the XR device. The information blending module 316 can perform this blending operation, for example, to ensure that the representation of real-world objects is correctly occluded by the digital content generated by the XR device. The display lens correction module 318 can process the generated virtual view to correct the geometric distortion and chromatic aberration of one or more display lenses used by the XR device to display the virtual view. For example, the display lens correction module 318 can use calibration data previously provided to the XR device or otherwise obtained by the XR device to pre-distort the virtual view before the XR device presents it.
[0065] Although Figure 3 illustrates an example of the pipeline 300 for depth-varying reprojection penetration in VST XR, various changes can be made to Figure 3 it. For example, Figure 3 the various components or functions in Figure 3 can be combined, further subdivided, replicated, omitted, or rearranged, and additional components or functions can be added according to specific requirements. As a specific example of this, the data fusion module 306 and the depth densification module 308 can be combined and implemented using a trained machine learning model (such as a trained convolutional neural network (CNN)). One or more suitable training datasets can be used to train the machine learning model, and these datasets can be used to train the machine learning model to process inputs such as sparse depth and low-resolution depth maps in order to generate dense depth maps. Additionally, Figure 3 the specific examples of the scene and the objects within the scene shown in
[0066] Figure 4 are for illustration and explanation only.
[0066] Figure 4 FIG. shows a specific example implementation of the pipeline 400 for depth-varying reprojection penetration in VST XR according to the present disclosure. The pipeline 400 herein can represent Figure 3 a more specific implementation of the pipeline 300 in Figure 4 . For ease of illustration, Figure 1 the pipeline 400 in Figure 2At least a portion of process 200. However, pipeline 400 can also be implemented using any other suitable device in any other suitable system, and pipeline 400 can perform any other suitable process.
[0067] As Figure 4 shown, sensor module 302 can include various sensors for capturing data for depth change reprojection. In this example, sensor module 302 includes a stereo camera pair 402 that can be used to capture stereo images of a scene. Sensor module 302 also includes at least one depth sensor 404, such as a time-of-flight (ToF) sensor, that can be used to measure depth within the scene. Sensor module 302 also includes a pose tracking stereo camera pair 406 and a position tracking sensor IMU 408 that can be used to identify the position and orientation of the XR device within the environment.
[0068] Data capture module 304 can include various functions for capturing or otherwise obtaining data using the various sensors of sensor module 302. In this example, data capture module 304 includes a perspective image capture function 410 for capturing perspective image pairs generated using camera 402. Data capture module 304 also includes a depth capture function 412 for capturing depth maps or other depth information generated using depth sensor 404. Data capture module 304 also includes a pose tracking image capture function 414 for capturing images generated using camera 406. Additionally, data capture module 304 includes an IMU data capture function 416 for capturing information generated using IMU 408.
[0069] The data fusion module 306 may include: a sparse depth calculation function 418 for generating a sparse depth map or other sparse depth values based on a stereo image pair of a scene. For example, the sparse depth calculation function 418 may calculate sparse depth values based on the stereo image pair using Structure from Motion (SFM) techniques. In other words, the sparse depth calculation function 418 may use the differences between the stereo image pair to estimate the 3D structure from a two-dimensional (2D) image and estimate the depth of different points in the 2D image based on the 3D structure. A sensor depth-to-resolution mapping function 420 is used to map and transform the sparse depth map or other sparse depth values from the depth sensor 404, which may involve processing the depth map or other depth values to have the same resolution as the perspective image captured using the camera 402. For example, the mapping function 420 may apply interpolation or other functions to the values in the sparse depth map from the depth sensor 404 to increase the resolution of the sparse depth map. A depth fusion function 422 may fuse the sparse depth points and the mapped depth map together to create a fused depth map associated with the stereo image from the camera 402. The fused depth map may still represent a sparse depth map and may, in some cases, have the same resolution as the perspective image from the camera 402.
[0070] The depth densification module 308 is used to densify or increase the number of depth values in the fused depth map to generate a dense depth map. The depth densification module 308 may also perform super-resolution processing on the fused depth map to improve the overall resolution of the dense depth map. The depth densification module 308 may include: a weight calculation function 424 for determining the weights to be applied to the depth values in the fused depth map. The weights may be determined in any suitable manner (e.g., by using the perspective image, camera pose, and spatial (position) information). This may be implemented to assign higher weights to some depth estimates and lower weights to other depth estimates. A minimization function 426 uses the weights and depth values in the fused depth map to minimize some criterion function, which may involve minimizing the value obtained from the criterion function calculated using the depth and weights. A dense depth map generation function 428 uses the value determined by the minimization function 426 to generate a dense depth map, which may include more depth values compared to the original sparse depth map.
[0071] Figure 4 An optional user focus tracking module 430 is shown, and in some embodiments, the user focus tracking module 430 may be used to identify the location within the scene where the user of the XR device appears to be focusing. It should be noted that although Figure 3This module is not shown, but the user focus tracking module 430 can be used in the pipeline 300. The user focus tracking module 430 can include: a user focus tracking and extraction function 432 for performing eye tracking and eye gaze estimation. For example, the user focus tracking and extraction function 432 can estimate the position where the user appears to be gazing and the coordinates of the position where the user appears to be gazing. The user focus area extraction function 434 can identify the approximate area where the user appears to be gazing based on the coordinates of the position where the user appears to be gazing. In some cases, for example, the user focus area extraction function 434 can identify the user's focus area as the area within a certain distance from the coordinates of the position where the user appears to be gazing. The foreground and background definition function 436 can be used to identify the foreground area of each scene and the background area of each scene. In some embodiments, the foreground area of a scene can be defined as the user's focus area, and other areas can be defined as the background area of the scene. However, it should be noted that the foreground area and the background area of each scene can be defined in any suitable way, for example, by identifying the foreground object as being within the threshold distance of the XR device.
[0072] The library 310 can include various known 3D models for scene and object reconstruction. For example, the library 310 can include 3D models 438 of previously encountered and reconstructed scenes. The library 310 can also include 3D models 440 of previously encountered and reconstructed objects. The library 310 can also include 3D models 442 of virtual objects that were previously generated and may be inserted into other virtual views. It should be noted that the library 310 can selectively include 3D models of predefined scenes, objects, or virtual objects that may or may not have been previously encountered or generated by the XR device.
[0073] The scene model calculation and reconstruction module 312 can include: a noise reduction function 444 for reducing noise in the dense depth map, for example, by filtering the depth values in the dense depth map from the depth densification module 308. It should be noted that the noise reduction function 444 can use any suitable filtering or other noise reduction techniques. The volume-based 3D reconstruction function 446 can be used to reconstruct the scene based on the 3D model of the scene and the 3D models of the objects within the scene (possibly together with the 3D models of virtual objects). For example, the volume-based 3D reconstruction function 446 can determine the 3D models of the objects within the scene and how the 3D models of one or more virtual objects are positioned within the 3D model of the scene. This can enable the generation of 3D objects and scene models for each scene. It should be noted that the 3D models of the scene and the objects within the scene can be generated based on the dense depth map and / or retrieved from the library 310. The refinement function 448 can be used to refine each 3D object and scene model, for example, by refining the positions of the objects within the scene and providing any required object occlusion.
[0074] During the 3D model reconstruction, the library 310 can be searched to determine whether there is a 3D model of the same or similar scene or the same or similar object in the library 310. If a suitable matching 3D model of the scene or object is found in the library 310, the scene or object model can be directly used during the 3D model reconstruction (without recreating the model). If a suitable matching 3D model of the scene or object is not found in the library 310, the 3D model of the scene or object can be generated by the scene model calculation and reconstruction module 312 and stored in the library 310 for future use. In some embodiments, the library 310 may include 3D models of scene reconstruction with depth maps, 3D models of object reconstruction with depth maps, and 3D models of virtual objects with depth maps.
[0075] The virtual view generation module 314 may include: a scene and object model adjustment function 450, which can adjust the scene and object models for each scene to support rendering from a new viewpoint. For example, the model adjustment function 450 can transform and scale the 3D models obtained from the library 310 to adapt to the currently captured scene, and can use the scene and object models to determine the way a specific scene and the objects within that specific scene are presented from a specific viewpoint. For example, the specific viewpoint may represent the positions of the user's left and right eyes. The virtual view generation function 452 can generate virtual views from the new viewpoint using the scene and object models. For example, the virtual view generation function 452 can perform depth-based reprojection of the 3D models of the objects in the scene foreground to the left and right virtual views, and perform constant-depth reprojection of the 3D models of the objects in the scene background to the left and right virtual views. The virtual view refinement function 454 can refine the generated virtual views, for example, by performing ray tracing and lighting operations, to improve the quality of the virtual views.
[0076] The information mixing module 316 may include: a virtual object creation and selection function 456, which can be used to select one or more 3D models for one or more virtual objects to be inserted into the left and right virtual views (if needed). For example, the creation and selection function 456 can identify one or more virtual objects to be inserted into the left and right virtual views and determine whether there is already a 3D model 442 suitable for the one or more virtual objects in the library 310. If so, the creation and selection function 456 can retrieve the 3D model 442 of the virtual object from the library 310. If not, the creation and selection function 456 can create one or more 3D models for one or more new virtual objects. As described above, any new 3D model can be stored in the library 310 for future use. The creation and selection function 456 can use any suitable criteria or guidelines to determine whether to insert any virtual objects into the left and right virtual views and select one or more virtual objects.
[0077] The parallax correction function 458 can be used to perform parallax correction on one or more virtual objects to be inserted into the virtual view. The parallax correction can be based on the real-world depth in the scene and take into account the fact that each virtual object may appear at different positions in the scene for the user's left and right eyes depending on the apparent depth of the virtual object within the scene. The parallax correction function 458 can use any suitable technique to perform the parallax correction. The virtual object blending and overlapping function 460 can blend one or more virtual objects into the captured scene represented by the associated left and right virtual views. In some cases, the virtual object blending and overlapping function 460 can place at least one virtual object on top of another virtual object or on top of at least one real-world object within the left and right virtual views before blending. The virtual object blending and overlapping function 460 can use any suitable technique to blend the image data.
[0078] The display lens correction module 318 can include: a lens calibration and distortion model generation function 462, which can be used to calibrate one or more lenses or other components of at least one display of the XR device and generate a model that represents the way an image is distorted by the one or more lenses. For example, this may involve determining the position of one or more lenses or other components of the display of the XR device and determining the way the lens distorts the image passing through the lens at that position. In some cases, at least some of such information can be predefined and stored on the XR device. The geometric distortion correction function 464 and the chromatic aberration correction function 466 can be used to preprocess the generated virtual view (modified by the information blending module 316) in order to pre-compensate for the virtual view distortion and aberration caused by one or more lenses or other components of at least one display of the XR device. For example, the geometric distortion correction function 464 can pre-distort the virtual view in order to pre-compensate for the warping or other spatial distortion caused by the lens; and the chromatic aberration correction function 466 can pre-distort the virtual view in order to pre-compensate for the chromatic aberration caused by the lens. The distortion center and field of view (FOV) calibration function 468 can be used to calibrate the distortion center and field of view of one or more lenses, for example, by modifying the position and aiming direction of the lens.
[0079] At this time, the left virtual view and the right virtual view can be rendered and presented on one or more displays of the XR device. For example, the XR device can display the left virtual image and the corresponding right virtual image to the user on one or more displays 160 of the electronic device 101, for example. This can be repeated any number of times to provide the user with a sequence of left virtual images and right virtual images. It should be noted that the displays used herein can represent two separate display panels (e.g., a left display panel and a right display panel that can be viewed separately by the user's eyes), or can be a single display panel (e.g., the left and right portions of the display panel that can be viewed separately by the user's eyes).
[0080] Although Figure 4 shows a specific example implementation of the pipeline 400 for depth change reprojection penetration in VST XR, various changes can be made to Figure 4 it. For example, Figure 4 the various components or functions in it can be combined, further subdivided, replicated, omitted, or rearranged, and additional components or functions can be added according to specific requirements. In addition, each module in the pipeline 400 can be implemented in any other suitable manner using the functions required to perform the operations of that module.
[0081] Figure 5 shows an example process 500 for 3D model reconstruction using image-guided depth fusion according to the present disclosure. For example, Figure 5 the process 500 shown can be performed using Figure 3 or Figure 4 the data fusion module 306 and the scene model calculation and reconstruction module 312. As Figure 5 shown, the process 500 receives a perspective stereo image pair 502 from the camera 402, and the perspective stereo image pair 502 can be provided by the perspective image capture function 410. The process 500 also receives a depth map 504 or other depth information, and the depth map 504 or other depth information can be based on data from the depth sensor 404 provided by the depth capture function 412. The process 500 further receives perspective camera pose information 506, and the perspective camera pose information 506 can be based on data from the camera 406 and the attitude tracking sensor IMU 408 provided by the attitude tracking image capture function 414 and the IMU data capture function 416. In addition, the process 500 receives a sparse depth map or other sparse depth 508, and the sparse depth map or other sparse depth 508 can be generated by the sparse depth calculation function 418.
[0082] In this example, the depth map from the depth fusion function 510 is used to process the depth map 504, the pose information 506, and the sparse depth 508. The depth map from the depth fusion function 510 can be performed, for example, as part of the depth fusion function 422 and the noise reduction function 444 described above. The depth fusion function 510 can combine the depth values determined using the depth map 504, the pose information 506, and the sparse depth 508 to generate a depth map for each perspective stereo image pair 502. Ideally, the depth map generated by the depth fusion function 510 includes only the true depth 512 (without noise). However, in practice, each depth map generated by the depth fusion function 510 includes the true depth 512 affected by one or more noise sources 514. Various noise sources 514 can affect the accuracy of the depth map generated by the depth fusion function 510, such as depth sensor noise affecting the depth map 504, pose sensor noise affecting the pose information 506, and other depth noise.
[0083] To help compensate for the various noise sources 514, the process 500 uses the stereo image pair consistency determination function 516 to identify the consistency or inconsistency between each stereo perspective image pair 502 being processed. It is known that a stereo image pair may include the same points located at different positions, depending on the depth of these points within the imaged scene. Therefore, the stereo image pair consistency determination function 516 can use the consistency of each perspective stereo image pair 502 and other information (such as various inputs of the depth map from the depth fusion function 510) to verify the correctness of the depth generated by the depth map of the depth fusion function 510 and detect noise in the generated depth map.
[0084] Figure 6 An example consistency between the left and right images in a stereo image pair according to the present disclosure is shown. As Figure 6 shown, the left image 600 and the right image 602 form a stereo image pair and can capture an object (in this case a tree) in the scene. If a point 606 on an object at a specified coordinate in the 3D space is represented as , then this point 606 is associated with a point 608 in the left image 600 and a point 610 in the right image 602. The point 608 can be represented as , which represents the projection of the point 606 at the coordinate on the image plane of the left image 600. Similarly, the point 610 can be represented as , which represents the projection of the point 606 at the coordinate on the image plane of the right image 602. Assuming that the images 600 and 602 have been rectified such that their corresponding epipolar lines are collinear and parallel to the axis, there is only a parallax in the direction between the two points 608 and 610.
[0085] The relationship between the left image 600 and the right image 602 can be expressed as follows.
[0086] (1)
[0087] In this article, represents the focal lengths of the left and right cameras or other imaging sensors, represents the distance between the left and right cameras, and represents the depth of the point . Additionally, represents the coordinates of point 608, and represents the coordinates of point 610. In Figure 6 , the terms and represent the positions where the line parallel to the axis and passing through the centers of images 600 and 602 intersects the axis (which may also be the position where the user's eyes are expected to be located). The term " " in formula (1) represents the parallax between two points 608 and 610 in images 600 and 602. Based on this, the coordinates and can be associated with the 3D coordinates in the following manner, for example.
[0088] (2)
[0089] (3)
[0090] As can be seen from this article, the depth of point 606 is related to the positions and There is a relationship, which is related to the consistency between the left image 600 and the right image 602 of the stereoscopic image pair. More specifically, when the point 606 is closer to the camera (meaning the depth d is smaller), the parallax between the points 608 and 610 is larger. When the point 606 is farther from the camera (meaning the depth d is larger), the parallax between the points 608 and 610 is smaller. Therefore, it is possible to use the parallax between the respective points 608 and 610 in the perspective stereoscopic image 502 to verify whether the depth determined by the depth map generated by the depth fusion function 510 using other inputs looks correct. If the decision function 518 detects that the depth determined by the depth map from the depth fusion function 510 roughly matches the depth determined using the consistency between the perspective stereoscopic image 502 (e.g., within a threshold amount or percentage range), then little or no modification may be required to the depth determined by the depth map from the depth fusion function 510. If the depth determined by the depth map from the depth fusion function 510 does not roughly match the depth determined using the consistency between the perspective stereoscopic image 502, then the depth determined by the depth map from the depth fusion function 510 can be modified, for example, by performing some type of noise reduction processing on the depth map from the depth fusion function 510. Any suitable noise reduction technique can be used to help make the depth generated by the depth map from the depth fusion function 510 generally consistent with the depth determined using the consistency between the perspective stereoscopic image 502.
[0091] This generates a noise-reduced depth map 520 for each perspective stereoscopic image pair 502. Each noise-reduced depth map 520 can be processed using the volume reconstruction function 522. The volume reconstruction function 522 is used to convert the depth data in the noise-reduced depth map 520 into a 3D model of the captured scene and the objects within that scene. The volume reconstruction function 522 can use any suitable technique to generate the 3D model of the scene and the objects, such as truncated signed distance field (TSDF) volume reconstruction. The color texture extraction function 524 can process the perspective stereoscopic image 502 to extract texture information of different colors in the perspective stereoscopic image 502, such as the red, green, and blue image data of the perspective stereoscopic image 502. The extracted color texture is combined with the 3D model generated by the volume reconstruction function 522 to generate a 3D model 526, which is used to represent the captured scene in the perspective stereoscopic image 502 and the objects within the scene. The 3D model 525 can optionally be stored in the library 310 as the 3D model 438 of the scene and the 3D model 440 of the object.
[0092] Although Figure 5 shows an example of the 3D model reconstruction process 500 with image-guided depth fusion, various changes can be made to Figure 5 it. For example, Figure 5The various functions in can be combined, further subdivided, replicated, omitted, or rearranged, and additional components or functions can be added according to specific requirements. Although Figure 6 shows an example of the consistency between the left and right images of a stereoscopic image pair, various changes can be made to Figure 6 . For example, Figure 6 's content is only used to illustrate the way in which the depth of points can affect the positions of these points in the captured images. Figure 6 The specific examples shown in do not limit the scope of the present disclosure to any particular scenario or scenario content.
[0093] As described above, there are multiple ways to segment a scene into a foreground region with one or more foreground objects and a background region with one or more background objects. Example methods for performing this segmentation include segmentation based on user focus and segmentation based on distance. An example implementation of each of these methods will be described separately below.
[0094] Figure 7 shows an example process 700 for separating a scene into foreground and background objects based on user focus according to the present disclosure. Figure 7 The shown process 700 can be performed, for example, using Figure 3 or Figure 4 's data capture module 304 and Figure 4 's user focus tracking module 430. For ease of explanation, Figure 7 's process 700 is described as being executed or implemented by the electronic device 101 in Figure 1 's network configuration 100. However, process 700 can also be implemented using any other suitable device and in any other suitable system.
[0095] As Figure 7 shown, the data capture module 304 can include: a perspective image capture function 410 for capturing a pair of perspective images generated using the camera 402. The data capture module 304 can also include: an IMU data capture function 416 for capturing information generated using the IMU 408. The data capture module 304 can also include: a motion data capture function 702 for capturing information generated using one or more motion sensors 180, where one or more motion sensors 180 can sense the motion of the XR device. In addition, the data capture module 304 can also include: an eye / head tracking image capture function 704, which can capture images associated with the user's eyes or head.
[0096] The information captured using the IMU data capture function 416, the motion data capture function 702, and the eye / head tracking image capture function 704 is provided to the eye movement and eye gaze tracking operation 706. The tracking operation 706 operates to estimate the position that the user of the XR device appears to be looking (gazing) at. The tracking operation 706 may include: an eye detection and extraction function 708 that can identify and isolate the eyes of the user's face (e.g., in an image of at least a portion of the user's face). An eye movement tracking function 710 can use the isolated images of the user's eyes to estimate the movement of the user's eyes over time. For example, the eye movement tracking function 710 can determine whether the user appears to be gazing in the same direction over time or is switching his or her focus. A head pose estimation and prediction function 712 can process this information to estimate / predict the accurate pose of the user's head. An eye gaze estimation and tracking function 714 can use the pose of the user's head to estimate the position that the user appears to be gazing at and track the position that the user appears to be gazing at over time. An eye position and gaze coordinate identification function 716 uses the results of the eye movement tracking and the gaze estimation and tracking to identify the position and coordinates in the scene where the user appears to be gazing. It should be noted that the eye movement and eye gaze tracking operation 706 can use any suitable technique to identify and track the position that the user appears to be gazing at.
[0097] The perspective stereo image 502 and the estimated user gaze position are provided to a focus area determination operation 718, and the focus area determination operation 718 can identify the area where the user is focused based on the tracking. The focus area determination operation 718 may include a transformation function 720 that can transform the eye positions and coordinates determined by the tracking operation 706 into a different coordinate system. For example, the transformation function 720 can transform the eye positions and coordinates determined by the tracking operation 706 from the coordinate system used by the tracking operation 706 to the coordinates in the perspective camera coordinate system used by the camera 402. A user focus area determination function 722 can process the transformed eye positions and coordinates and the perspective image 502 to dynamically determine the user focus area within the scene captured in the perspective image 502. For example, the user focus area determination function 722 can identify the position in the perspective image 502 where the user appears to be gazing and define the area around that position, e.g., the area within a specified distance of the identified position.
[0098] The foreground / background region definition function 724 can use the user's focused region to define the foreground and background of the perspective image 502. For example, the foreground / background region definition function 724 can consider the user's focused region of each perspective image 502 as the foreground of the perspective image 502, and the foreground / background region definition function 724 can consider all other regions of each perspective image 502 as the background of the perspective image 502. At this time, the process 200 can perform depth-based reprojection on one or more 3D models associated with one or more objects within the foreground of the perspective image 502, and the process 200 can perform constant-depth reprojection on one or more 3D models associated with one or more objects within the background of the perspective image 502. This can be repeated over the entire sequence of perspective image pairs to account for changes in the user's focused region.
[0099] It should be noted that if the user's focused region cannot be determined in this document, all objects in the perspective image 502 can be regarded as background objects. In this case, the 3D models of all objects captured in the perspective image 502 can be subjected to constant-depth reprojection to generate a virtual view for presentation by the XR device. If the user once focuses on a specific region, the process 700 can update the virtual view by performing depth-based reprojection on any object currently located within the user's focused region.
[0100] Figure 8 An example process 800 of separating a scene into foreground objects and background objects using machine learning according to the present disclosure is shown. Figure 8 The illustrated process 800 can be performed, for example, using Figure 3 or Figure 4 the depth densification module 308, the scene model calculation and reconstruction module 312, and the virtual view generation module 314 in Figure 8 For the sake of illustration, Figure 1 the process 800 in
[0101] As Figure 8 shown, the process 800 receives a sequence of perspective images 802, which can be captured using the camera 402. The process 800 also receives a sequence of depth maps 804, which can be captured using at least one depth sensor 404, or generated in any other suitable manner. The sequence of depth maps 804 corresponds to the sequence of perspective images 802.
[0102] A perspective image sequence 802 is provided to a trained machine learning model 806, which has been trained to perform object detection and extraction. The machine learning model 806 can include or support an object detection function 808 and a feature extraction function 810. The object detection function 808 can operate to detect specific objects contained in the perspective image sequence 802, and the feature extraction function 810 can be trained to identify specific features of the objects detected in the perspective image sequence 802. This results in the identification of the detected objects 812 and the extracted features 814. The machine learning model 806 represents any suitable machine learning-based architecture that has been trained to perform object recognition and feature extraction, such as a deep neural network (DNN) or other neural network. As a specific example, the machine learning model 806 can represent a DNN with a multi-scale architecture and a two-dimensional convolutional network. After training (e.g., using a labeled dataset), the machine learning model 806 can detect objects of different sizes in the captured scene. The DNN can also include convolutional layers for extracting multi-scale image features of the image regions containing the detected objects.
[0103] A depth map post-processing operation 816 processes the depth map sequence 804 to generate a dense depth map 818. For example, the depth map post-processing operation 816 can include a depth denoising function 820 and a depth densification function 822. These functions 820 and 822 can be implemented in the same or a similar manner as the corresponding functions described above. In some cases, the depth map post-processing operation 816 can be implemented using a trained machine learning model that has been trained to reduce or eliminate noise, such as noise introduced by depth capture and depth reconstruction. The machine learning model can also be trained to perform hole filling to fill in missing depth values. As needed, the machine learning model can also perform super-resolution processing to upscale the depth map to the same resolution as the perspective image. The machine learning model represents any suitable machine learning-based architecture that has been trained to perform denoising and / or depth densification, such as a DNN or other neural network.
[0104] The background segmentation function 824 processes the perspective image sequence 802, the detected objects 812, the extracted features 814, and the dense depth map 818 to identify the background objects in the perspective image. For example, the background segmentation function 824 can generate and output an image containing the background objects in the perspective image. In some cases, the background segmentation function 824 can process patches of the detected objects 812, the extracted features 814 of the detected objects 812, and the dense depth map 818 of the detected objects. In some embodiments, the background segmentation function 824 can be implemented using a trained machine learning model (e.g., a model with a trained encoder-decoder network 826). Additionally, in some embodiments, an L2 loss criterion can be used to train the trained machine learning model. During training, the machine learning model can learn the background image information contained in the training images and learn ways to generate images containing the identified background image information.
[0105] The foreground object segmentation function 828 processes the perspective image sequence 802, the dense depth map 818, and the background image generated by the background segmentation function 824 to identify the foreground objects in the perspective image. For example, the foreground object segmentation function 828 can be used to segment the foreground objects from the perspective image sequence 802. In some embodiments, the foreground object segmentation function 828 can be implemented using a trained machine learning model (e.g., a model with a trained fully convolutional network 830). Additionally, in some embodiments, a cross-entropy loss criterion can be used to train the trained machine learning model. During training, the machine learning model can learn to identify the foreground objects contained in the training images and learn ways to segment the foreground objects from the training images. The depth determination function 832 can use the results of the foreground object segmentation function 828 (and possibly the background segmentation function 824) and can automatically generate depth thresholds that can be used to separate the foreground objects from the background image.
[0106] For the foreground object segmentation function 828, the foreground object segmentation problem can be regarded as a "binary classification" problem. That is, the foreground object segmentation function 828 can classify each pixel in the perspective image into one of two categories: the first category if the pixel is part of the foreground object; and the second category if the pixel is not part of the foreground object. The fully convolutional network 830 can be trained and classified using a cross-entropy loss criterion (also known as a loss function). The cross-entropy loss can be used when adjusting the model weights during training. A smaller loss indicates that the model is trained better, and a cross-entropy loss of zero indicates a perfect model. Since the fully convolutional network 830 uses two categories, a binary cross-entropy loss function can be used. In some cases, this loss function can be expressed as follows.
[0107] (4)
[0108] In this text, represents a binary loss, represents the ground truth, and represents the softmax probability of a class with the ground truth In some cases, the softmax probability can be expressed as follows.
[0109] (5)
[0110] In this text, represents the softmax probability of a class and represents the logit of a class In some cases, the average cross-entropy over all data examples used during training can be determined. In some cases, the average cross-entropy can be expressed as follows.
[0111] (6)
[0112] In this text, represents the number of data points.
[0113] In some embodiments, supervised learning can be used to train one or more machine learning models used in process 800. Additionally, in some embodiments, training data for different locations or environment types can be used to train one or more machine learning models used in process 800. Example locations or environment types can include kitchens, living rooms, bedrooms, offices, lobbies, and other locations. Universal identifiers can be provided for objects in various locations or environments, such as desktops, walls, cups, keyboards, mice, medicine bottles, etc. This can allow one or more machine learning models to provide different threshold distances for segmenting foreground objects and background objects in different environments. Ideally, one or more machine learning models used in process 800 are well-trained to determine near (foreground) and far (background) designations for new objects in new scenarios.
[0114] It should be noted that the operation of process 800 may vary based on other factors unrelated to the actual perspective image content. For example, if the XR device experiences high computational load or high heat generation, or if the XR device needs to conserve power, then process 800 can be configured to segment more objects into the background, for example, based on the likelihood of each object being an object of user interaction and / or user interest. Since background objects can be reprojected with a constant depth rather than a depth-based reprojection, this helps reduce the computational load and / or power consumption of the XR device. As another example, one or more machine learning models used in process 800 can learn the relationship between the distance threshold and the experience, preferences, and / or interests of a particular user.
[0115] Although Figure 7 an example of process 700 for separating a scene into foreground and background objects based on user focus is shown, and Figure 8 an example of process 800 for separating a scene into foreground and background objects using machine learning is shown, various changes can be made to Figure 7 and Figure 8 them. For example, various components or functions in Figure 7 and FIG. 8 can be combined, further subdivided, replicated, omitted, or rearranged, and additional components or functions can be added according to specific needs. Additionally, although FIGS. 7 and 8 show two example techniques for separating a scene into foreground and background objects, the scene can also be segmented in any other suitable way. For example, a default threshold distance can be used instead of having one or more machine learning models identify a dynamic threshold distance. In some cases, the default threshold distance can be based on the user's long-term experience. As another example, the user can manually set the threshold distance, for example, based on the user's own experience, needs, or preferences.
[0116] Figure 9 An example depth-based reprojection 900 for generating left and right virtual views according to the present disclosure is shown. For example, when reprojecting the 3D model of an object determined to be in the foreground of a scene, the virtual view generation function 452 can perform the depth-based reprojection 900. As described above, the captured scene and the 3D models of the objects in the scene can be obtained using various operations in process 200 or pipeline 300 or 400. One or more 3D models can be obtained from library 310 if available.
[0117] As Figure 9 shown, the depth-based reprojection 900 can be performed to reproject the 3D model 902 of an object (defined as a cube in this example) onto the left virtual image frame 904 and the right virtual image frame 906. The left virtual image frame 904 is used to create a virtual view for the user's left eye 908, and the right virtual image frame 906 is used to create a virtual view for the user's right eye 910. Here, it is assumed that the virtual camera positions are at the positions of the user's left eye 908 and the user's right eye 910. Ideally, the created virtual views would appear as if captured by cameras at the user's eye positions, even though these cameras are actually located elsewhere. In this example, the ipd (interpupillary distance) is used to represent the separation of the virtual camera positions.
[0118] Based on the above formula (2), the points of the 3D model 902 can be projected onto the left virtual image frame 904 and the right virtual image frame 906. For example, assuming is a point on the 3D model 902, is the projection on the left virtual frame 904, and is the projection on the right virtual frame 906. Based on Equation (2), the following relationship can be obtained.
[0119] (7)
[0120] Therefore, this relationship can be used to re-project this point of the 3D model 902 onto the left virtual image frame 904 and the right virtual image frame 906. Repeating this projection for multiple points of the 3D model 902 can re-project the 3D model 902 onto each of the left virtual image frame 904 and the right virtual image frame 906.
[0121] It should be noted that this depth-based re-projection may occur only for 3D models associated with objects identified as being in the foreground of the scene (defined by the user's focus, threshold distance, or other parameters). One advantage of this method is that depth-based re-projection (i.e., re-projecting an object at different depth levels to generate virtual view images) can be limited to the objects that the user is most likely to be interested in. Other objects can be considered background objects, in which case, a simpler constant-depth re-projection can be performed on the 3D models of those objects.
[0122] FIG. 10 shows an example depth-varying re-projection 1000 for objects selected at different depth levels according to the present disclosure. For example, when re-projecting the 3D models of multiple objects at different depths within a scene 1002, the virtual view generation function 452 can perform a depth-varying re-projection 1000. In this example, the scene 1002 includes three objects, and two threshold distances d1 and d2 are identified. The left virtual image frame 1004 and the right virtual image frame 1006 of the scene 1002 can be generated in the manner described above, for example.
[0123] If process 800 determines that the threshold distance d1 should be used, only the foremost object (i.e., the horse) in scene 1002 can undergo depth-based reprojection, while other objects can undergo constant-depth reprojection when generating the left virtual image frame 1004 and the right virtual image frame 1006. If process 800 determines that the threshold distance d2 should be used, only the two foremost objects (i.e., the horse and the tiger) in scene 1002 can undergo depth-based reprojection, while other objects can undergo constant-depth reprojection when generating the left virtual image frame 1004 and the right virtual image frame 1006. In some cases, constant-depth reprojection can be generated at the identified threshold distance d1 or d2.
[0124] This method of performing depth-based reprojection on some objects and constant-depth reprojection on other objects is very useful for video see-through XR applications. For example, assume that a user is using a physical keyboard with an XR headset. The user can type on the physical keyboard using their fingers, and the XR headset can identify the keys pressed by the user in order to obtain user information (even if the physical keyboard is not connected to the XR headset and may not even be operational or receiving power). The physical keyboard can be rendered in the virtual image as accurately as possible so that the XR headset can accurately identify the keys pressed by the user. Other objects can be treated as background objects, and constant-depth reprojection can be performed on their 3D models. This method is also very beneficial because precise depth reconstruction can be computationally expensive, so performing depth-based reprojection with precise depth on only one or some objects (rather than all objects) in the scene can reduce the computational complexity of the virtual view generation process.
[0125] It should be noted that there are some cases where at least a part of an object can be within the specified threshold distance (the threshold distance determined using process 800 or determined in any other suitable manner), while at least another part of the object can be outside the specified threshold distance. In these cases, the object can be considered a foreground object or a background object, which can depend on the default settings or user settings of the XR device, or some other criteria or guidelines. In some embodiments, since at least a part of the object is within the threshold distance, the object can be considered a foreground object. In a particular embodiment, if at least a part of at least one particular type of object is within the threshold distance, the object can be considered a foreground object; while if at least a part of at least one other type of object is outside the threshold distance, the object can be considered a background object.
[0126] Although FIG. 9 shows an example of depth-based reprojection 900 for generating left and right virtual views, and FIG. 10 shows an example of depth-varying reprojection 1000 for objects selected at different depth levels, various changes can be made to FIGS. 9 and 10. For example, the specific examples of the scenes and the objects within the scenes shown in FIGS. 9 and 10 are for illustration and explanation only.
[0127] As described above, the library 310 can be used to store 3D models of scenes, objects, and possibly virtual objects. At least some of the 3D models stored in the library 310 may include 3D models of scenes previously encountered by the XR device and 3D models of objects in those scenes. FIG. 11 shows an example process 1100 for constructing the library 310 of 3D reconstructed scenes and objects according to the present disclosure. As shown in FIG. 11, the process 1100 includes two general parts: a subprocess 1102 for reconstructing 3D models of scenes and objects, and a subprocess 1104 for constructing the library. It should be noted that the subprocess 1102 shown herein can actually be used as part of the scene model calculation and reconstruction module 312 described above. For example, when the scene model calculation and reconstruction module 312 uses the subprocess 1102 to generate 3D models of the scenes and the objects in those scenes captured in the perspective image 502.
[0128] The subprocess 1102 in this example can include: a key object detection and extraction function 1106, which can process the perspective image 502 to identify specific types of objects included in the perspective image 502 and extract those objects from the perspective image 502. In some cases, the detection and extraction function 1106 can be used to identify and extract predefined types of objects from an image of a scene. In other cases, for example, when the detection and extraction function 1106 is configured to identify objects based on the discrete or individual structures of the objects in the perspective image 502, the detection and extraction function 1106 can be more dynamic. A key object segmentation function 1108 can be used to remove the identified key objects from the perspective image 502, thereby leaving the background in the processed image. This results in the generation of a scene segmentation 1110 that includes only the background and any non-key objects contained in the perspective image 502, and an object segmentation 1112 that includes only the key objects extracted from the perspective image 502.
[0129] Scene segmentation 1110 is provided to a scene reconstruction function 1114, which can process the scene segmentation 1110 to generate a 3D model of the scene captured in the perspective image 502. The scene reconstruction function 1114 can use any suitable technique, such as structure from motion technique, to generate the 3D model of the scene captured in the perspective image 502 using the scene segmentation 1110. Object segmentation 1112 is provided to an object reconstruction function 1116, which can process the object segmentation 1112 to generate 3D models of multiple objects captured in the perspective image 502. The object reconstruction function 1116 can use any suitable technique, such as structure from motion technique, to generate the 3D models of multiple objects captured in the perspective image 502 using the object segmentation 1112.
[0130] The sub - process 1104 in this example can receive a 3D model 1118, which can include the 3D model of the scene generated by the scene reconstruction function 1114 and the 3D models of the objects generated by the object reconstruction function 1116. A library query function 1120 can be used to determine whether each 3D model 1118 is the same as or similar to any 3D model in the library 310. A determination function 1122 can process the query result and can store the 3D model 1118 in the library 310 or not store the 3D model 1118 in the library 310. For example, if the 3D model 1118 is new and has not been stored in the library 310 before, the 3D model 1118 can be stored in the library 310.
[0131] In some embodiments, each 3D model can be defined using the following types of data. A 3D mesh can be used to represent the geometry of the associated scene or object, e.g., when the 3D mesh describes the 3D shape of the associated scene or object. For example, a color texture map can be used to represent the color of the associated scene or object when it describes the color of the surface or other regions of the associated scene or object. One or more parameters can identify at least one of the pattern, size, pose, and transformation of the associated scene or object. For example, when reconstructing the 3D model, one or more of these parameters can be attached to the 3D model. With these parameters, it is possible to unify scenes or objects based on unit scenes or objects with a uniform size, uniform pose, and uniform transformation.
[0132] To use the content of library 310, the scene model calculation and reconstruction module 312 can use camera tracking to obtain camera pose and object detection information, which can be implemented as detecting objects in the captured scene. The scene model calculation and reconstruction module 312 can perform functions 1106, 1108, 1114, and 1116 to segment the captured scene into foreground and background, and can search for potentially matching 3D models in library 310 using the objects in the foreground (e.g., by using pattern matching). The pattern matching herein can look for 3D models with the same or similar 3D meshes, the same or similar color texture maps, and the same or similar object parameters (at least within a certain threshold amount or percentage range). If a match is found, the volume-based 3D reconstruction function 446 can use the 3D model from library 350. For example, a transformation can be determined based on the current camera pose, and this transformation can be used to transform the unit objects from library 310 into the captured scene, and then the calculated scaling factor and orientation can be used to fit the transformed objects to the positions in the scene. After all objects are processed (by reusing 3D models from library 310 and / or reconstructing new 3D models), the left virtual view and the right virtual view can be generated as described above, while using accurate depth-based reprojection for foreground objects and simpler constant-depth reprojection for background objects.
[0133] In some cases, the 3D models 1118 in library 310 can include predefined 3D models or other 3D models provided by an external source (e.g., the manufacturer of the XR device). Thus, library 310 can be built in both an online manner and an offline manner. In the online case, the 3D models generated by the XR device can be stored in library 310 because these 3D models are for scenes or objects not previously encountered by the XR device. In the offline case, images of 3D scenes can be captured and used to reconstruct 3D models of the scenes and objects (which may or may not appear on the XR device), and then these 3D models can be added to library 310.
[0134] Although FIG. 11 shows an example of the process 1100 for building library 310 for 3D reconstruction of scenes and objects, various changes can be made to FIG. 11. For example, the use of library 310 is optional, so the sub-process 1104 shown Figure 11 can be omitted.
[0135] In the above discussion, the XR device was described as capturing and processing perspective images so that, for example, these perspective images are used to generate virtual views by re-projecting 3D models of foreground and background objects. This can be implemented as being repeated (e.g., continuously) as the user moves the XR device in the environment. However, in some cases, it is possible to generate some virtual views in this way and then generate additional virtual views based on earlier virtual views (rather than based on image capture and processing operations).
[0136] FIG. 12 shows an example of generating a virtual image 1200 at one viewpoint using another virtual image located at another viewpoint according to the present disclosure. As shown in FIG. 12, the above-described technique can be used to generate a virtual image 1202 at a location herein referred to as "viewpoint A". Assume that the user moves and subsequently arrives at another location herein referred to as "viewpoint B". In some embodiments, the user's XR device can generate another virtual image 1204 at the other location using the same technique as described above.
[0137] However, in some embodiments, the user's XR device can perform a warping operation (e.g., 3D warping) to generate the virtual image 1204 using the virtual image 1202. That is, instead of obtaining a 3D model at the second viewpoint and performing re-projection to generate the virtual image 1204, the user's XR device can warp the virtual image 1202 to generate the virtual image 1204. In some cases, the 3D warping or other warping operations may still involve generating a dense depth map that can be used to support the warping operation. However, the generation of the virtual image 1204 may not involve any depth-based or constant-depth re-projection of a 3D model.
[0138] Although FIG. 12 shows an example of generating a virtual image 1200 at one viewpoint using another virtual image at another viewpoint, various changes can be made to FIG. 12. For example, the specific examples of the scene and the objects within the scene shown in FIG. 12 are for illustration and explanation only. In addition, for ease of illustration, the distance between the two viewpoints is exaggerated herein, and warping can be implemented for any suitable change in the viewpoint position of the XR device.
[0139] FIG. 13 shows an example method 1300 of re-projecting penetration with depth change in VST XR according to the present disclosure. For ease of illustration, the method 1300 shown in FIG. 13 is described as being performed by Figure 1is performed by the electronic device 101 in the network configuration 100, where the electronic device 101 can perform the process 200 shown in FIG. 2 using the pipeline 300 shown in FIG. 3 or the pipeline 400 shown in FIG. 4. However, the method 1300 shown in FIG. 13 can also be performed using any other suitable device and pipeline and in any other suitable system.
[0140] As shown in FIG. 13, in step 1302, a scene image is obtained using the stereoscopic imaging sensor of the XR device, and depth data associated with the image is obtained in step 1304. This can include, for example, the processor 120 of the electronic device 101 capturing a perspective image 502 of the scene using the camera 402 or other imaging sensor 180, where the scene includes a plurality of objects captured in the perspective image 502. This can also include the processor 120 of the electronic device 101 obtaining a depth map or other depth data using the depth sensor 404, although the depth data can be obtained in any other suitable manner. In some cases, the depth map or other depth data can be preprocessed, for example, by performing super-resolution processing and depth densification. It should be noted that any other input data, such as camera pose, IMU, head pose, or eye tracking data, can be received herein.
[0141] A volume-based 3D model of the objects in the scene is obtained in step 1306. This can include, for example, the processor 120 of the electronic device 101 searching the library 310 to determine whether any object in the scene is associated with a predefined 3D model or other 3D models stored in the library 310. If so, the processor 120 of the electronic device 101 can retrieve the 3D model or other 3D models from the library. If not, the processor 120 of the electronic device 101 can perform object reconstruction to generate one or more 3D models for one or more objects in the scene. In some cases, the processor 120 of the electronic device 101 can perform object reconstruction by performing volume reconstruction based on the dense depth map and generating a new 3D model representing the object based on the volume reconstruction and at least one color texture captured in the perspective image 502. Additionally, in some cases, this can include the processor 120 of the electronic device 101 obtaining a 3D model of the scene itself.
[0142] In step 1308, the scene is segmented into one or more foreground objects and one or more background objects. It should be noted that any other suitable segmentation technique can be used herein. For example, in some cases, this can include the processor 120 of the electronic device 101 identifying the focused area of the user of the electronic device 101 and identifying any objects within the focused area, where the identified objects can be regarded as one or more foreground objects (and all other objects can be regarded as background objects). In other cases, this can include the processor 120 of the electronic device 101 using a trained machine learning model to identify a threshold distance, where any objects within the threshold distance can be regarded as one or more foreground objects (and all other objects can be regarded as background objects).
[0143] In step 1310, for the foreground objects, perform depth-based reprojection of the 3D models of the foreground objects to the left virtual view and the right virtual view. For example, this can include the processor 120 of the electronic device 101 performing depth-based reprojection on the 3D models of each foreground object based on the depth of each foreground object within the scene. In step 1312, for the background objects, perform constant-depth reprojection of the 3D models of the background objects to the left virtual view and the right virtual view. For example, this can include the processor 120 of the electronic device 101 performing constant-depth reprojection of the 3D model of each background object to a specified depth.
[0144] In step 1314, render the left virtual view and the right virtual view for presentation, and in step 1316, present the rendered left virtual view and right virtual view on the XR device. This can include, for example, the processor 120 of the electronic device 101 generating left and right virtual view images based on the reprojection. This can also include the processor 120 of the electronic device 101 performing any required operations to correct geometric distortion and chromatic aberration. This can also include the processor 120 of the electronic device 101 initiating the display of the rendered left virtual view and right virtual view on one or more displays 160 of the electronic device 101.
[0145] Although FIG. 13 shows an example of the depth-varying reprojection penetration method 1300 in VST XR, various changes can be made to FIG. 13. For example, although FIG. 13 is shown as a series of steps, the individual steps therein can overlap, occur in parallel, occur in a different order, or occur any number of times (including zero times).
[0146] It should be noted that the functions shown or described in FIGS. 2 to 13 can be implemented in any suitable manner in the electronic devices 101, 102, 104, the server 106, or other devices. For example, in some embodiments, at least some of the functions shown or described in FIGS. 2 to 13 can be implemented or supported by one or more software applications or other software instructions executed by the processor 120 of the electronic devices 101, 102, 104, the server 106, or other devices. In some embodiments, at least some of the functions shown or described in FIGS. 2 to 13 can be implemented or supported using dedicated hardware components. Generally, the functions shown or described in FIGS. 2 to 13 can be executed using any suitable hardware or any suitable combination of hardware and software / firmware instructions. In addition, the functions shown or described in FIGS. 2 to 13 can be executed by a single device or multiple devices.
[0147] Although the present disclosure has been described in connection with exemplary embodiments, those skilled in the art may still be suggested to make various changes and modifications. The present disclosure is intended to cover such changes and modifications that fall within the scope of the appended claims.
[0148] In an embodiment, a method for extended reality (XR) may include: determining whether a corresponding volumetric 3D model for an object exists in a previously generated library of volumetric 3D models. The method may include: obtaining the corresponding volumetric 3D model from the library in response to determining that the corresponding volumetric 3D model for the object exists in the library. The method may further include: generating a new volumetric 3D model representing the object in response to determining that the corresponding volumetric 3D model for the object does not exist in the library.
[0149] In an embodiment, the method may include: generating a dense depth map based on an image of the scene, depth data, and position data associated with the electronic device. The method may include: performing volumetric reconstruction based on the dense depth map. The method may include: generating a new volumetric 3D model representing the object based on the volumetric reconstruction and at least one color texture captured in the image.
[0150] In an embodiment, the pre-generated library of volumetric 3D models is configured to store 3D models of scenes, real-world objects, and virtual objects. In an embodiment, each 3D model includes a 3D mesh representing the geometry of the associated scene or object, a color texture map representing at least one color of the associated scene or object, and one or more parameters for identifying at least one of a pattern, size, pose, and transformation of the associated scene or object.
[0151] In an embodiment, one or more first objects are located within a threshold distance from the electronic device. In an embodiment, one or more second objects are located outside the threshold distance from the electronic device. In an embodiment, the method further includes: using a trained machine learning model to identify the threshold distance.
[0152] In an embodiment, the method may include: identifying a focus area of the user. In an embodiment, one or more first objects are located within the focus area of the user. In an embodiment, one or more second objects are located outside the focus area of the user.
[0153] In an embodiment, one or more first objects include one or more foreground objects in the scene. In an embodiment, one or more second objects include one or more background objects in the scene.
Claims
1. A method for extended reality XR, comprising: Obtaining (i) an image of a captured scene using a stereoscopic imaging sensor of an electronic device, and (ii) depth data associated with the image, the scene including a plurality of objects; Obtaining a volume-based three-dimensional 3D model of the plurality of objects included in the scene; For one or more first objects among the plurality of objects, performing depth-based reprojection of one or more 3D models of the one or more first objects to a left virtual view and a right virtual view based on one or more depths of the one or more first objects; For one or more second objects among the plurality of objects, performing constant-depth reprojection of one or more 3D models of the one or more second objects to the left virtual view and the right virtual view based on a specified depth; And Rendering the left virtual view and the right virtual view for presentation by the electronic device.
2. The method according to claim 1, wherein, Obtaining a volume-based 3D model of the plurality of objects in the scene includes: for each object among the plurality of objects, Determining whether a corresponding volume-based 3D model for the object exists in a previously generated library of volume-based 3D models; and Performing one of the following operations: In response to determining that a corresponding volume-based 3D model for the object exists in the library, obtaining the corresponding volume-based 3D model from the library; or In response to determining that a corresponding volume-based 3D model for the object does not exist in the library, generating a new volume-based 3D model representing the object.
3. The method according to claim 2, wherein, Generating a new volume-based 3D model representing the object includes: Generating a dense depth map based on the image of the scene, the depth data, and position data associated with the electronic device; Performing volume reconstruction based on the dense depth map; and Generating a new volume-based 3D model representing the object based on the volume reconstruction and at least one color texture captured in the image.
4. The method according to claim 2 or 3, wherein: The previously generated library of volume-based 3D models is configured to store 3D models of scenes, real-world objects, and virtual objects; and Each 3D model includes: a 3D mesh representing the geometry of an associated scene or object; a color texture map representing at least one color of the associated scene or object; and one or more parameters for identifying at least one of a pattern, size, pose, and transformation of the associated scene or object.
5. The method according to any one of claims 1 to 4, wherein: The one or more first objects are located within a threshold distance from the electronic device; The one or more second objects are located outside the threshold distance from the electronic device; And The method further includes: using a trained machine learning model to identify the threshold distance.
6. The method according to any one of claims 1 to 5, wherein: The method further includes: identifying a focus area of the user; The one or more first objects are located within the focus area of the user; and Said one or more second objects are located outside the user's focus area.
7. The method according to any one of claims 1 to 6, wherein: Said one or more first objects include one or more foreground objects in the scene; and Said one or more second objects include one or more background objects in the scene.
8. An electronic device (101) for extended reality XR, comprising: At least one display (160); An imaging sensor (180) configured to capture an image of a scene, the scene including a plurality of objects; And At least one processor (120) configured to: Obtain (i) the image of the scene and (ii) depth data associated with the image; Obtain a volume-based three-dimensional 3D model of the plurality of objects included in the scene; For one or more first objects among the plurality of objects, perform depth-based reprojection of the one or more 3D models of the one or more first objects to a left virtual view and a right virtual view based on the one or more depths of the one or more first objects; For one or more second objects among the plurality of objects, perform constant-depth reprojection of the one or more 3D models of the one or more second objects to the left virtual view and the right virtual view based on a specified depth; And Render the left virtual view and the right virtual view for presentation by the at least one display (160).
9. The electronic device (101) according to claim 8, wherein, To obtain a volume-based 3D model of the plurality of objects in the scene, the at least one processor (120) is configured for each object among the plurality of objects: Determine whether a corresponding volume-based 3D model for the object exists in a previously generated library of volume-based 3D models; and Perform one of the following operations: In response to determining that a corresponding volume-based 3D model for the object exists in the library, obtain the corresponding volume-based 3D model from the library; or In response to determining that a corresponding volume-based 3D model for the object does not exist in the library, generate a new volume-based 3D model representing the object.
10. The electronic device (101) according to claim 9, wherein, To generate a new volume-based 3D model representing the object, the at least one processor (120) is configured to: Generate a dense depth map based on the image of the scene, the depth data, and position data associated with the electronic device (101); Perform volume reconstruction based on the dense depth map; And Generate a new volume-based 3D model representing the object based on the volume reconstruction and at least one color texture captured in the image.
11. The electronic device (101) according to claim 8 or 9, wherein: The previously generated library of volume-based 3D models is configured to store 3D models of scenes, real-world objects, and virtual objects; and Each 3D model includes: a 3D mesh representing the geometry of an associated scene or object; a color texture map representing at least one color of the associated scene or object; and one or more parameters for identifying at least one of a pattern, size, pose, and transformation of the associated scene or object.
12. The electronic device (101) according to any one of claims 8 to 11, wherein: The one or more first objects are located within a threshold distance from the electronic device (101); The one or more second objects are located outside the threshold distance from the electronic device (101); and The at least one processor (120) is further configured to identify the threshold distance using a trained machine learning model.
13. The electronic device (101) according to any one of claims 8 to 12, wherein: The at least one processor (120) is further configured to identify a focus area of a user; The one or more first objects are located within the focus area of the user; And The one or more second objects are located outside the focus area of the user.
14. The electronic device (101) according to any one of claims 8 to 13, wherein: The one or more first objects include one or more foreground objects in the scene; and The one or more second objects include one or more background objects in the scene.
15. A machine-readable medium comprising instructions that, when executed, cause at least one processor to perform the following operations: Obtain (i) an image of a captured scene using a stereoscopic imaging sensor of an electronic device; and (ii) depth data associated with the image, the scene including a plurality of objects; Obtain a volume-based three-dimensional 3D model of the plurality of objects included in the scene; For one or more first objects of the plurality of objects, perform depth-based reprojection of one or more 3D models of the one or more first objects to a left virtual view and a right virtual view based on one or more depths of the one or more first objects; For one or more second objects of the plurality of objects, perform constant-depth reprojection of one or more 3D models of the one or more second objects to the left virtual view and the right virtual view based on a specified depth; And Render the left virtual view and the right virtual view for presentation by the electronic device (101).